Local LLMs often produce incorrect API route definitions despite syntactically valid output
Local LLMs generate API routes with incorrect imports and parameter names.
When local LLMs draft FastAPI route definitions, they frequently embed deprecated imports such as RequestValidationError from fastapi.exceptions, a module removed in version 0.110, and they also reference List[User] without pulling it from typing. Even when documentation is pinned, the models still output request_body where the correct syntax is Body(...).
Attempting to steer the output has yielded mixed results. Adding version‑specific documentation helped the basic import statements, yet the model continued to use inconsistent naming, for example preferring request_body inside Body(...) rather than the expected body.
Few‑shot prompting cut down on outright hallucinations, but it made the model over‑fit to repetitive CRUD patterns, causing it to repeat those operations even when a webhook handler was requested.
Integrating retrieval‑augmented generation with ChromaDB and LangChain raised accuracy by pulling real code snippets, though the 4,000‑token context window filled quickly, leaving roughly 2,000 tokens for reasoning after the retrieved chunks consumed most of the buffer.
A post‑generation linting stage that runs Ruff and performs AST parsing lifted the first‑pass compilation success rate to 85%, but each usable snippet now needs 12‑18 seconds of processing time and usually demands more than three rounds of correction.
Balancing model size with reliability suggests that fine‑tuning smaller models on framework patterns may be more effective than relying on the retry loops introduced by linting.
Exploring guidance mechanisms or LMQL‑style constraints could enforce valid imports and reduce the need for extensive post‑processing.
Assessing whether a larger 13B quantized model (q4_k_m) improves reasoning is still open, but it brings higher VRAM usage and slower inference speeds.
All Replies (4)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Switching from 7B to 13B stopped the hallucinations for me. Which version are you using? I spent the weekend battling a 7B model running via Ollama, aiming to generate valid FastAPI route handlers. Every third response invents a RequestValidationError import from fastapi.exceptions that was deprecated two versions ago, or suggests response_model=List[User] without importing List from typing. The model confidently outputs code that looks syntactically correct but fails at import time. I tried a few approaches: ## Can Version-Pinned Documentation Improve Code Generation? 1. Added a system prompt with version-pinned docs — I pasted the FastAPI 0.110 reference into the context window. This helped with imports, yet the model still invents parameter names like request_body instead of body for Body(...). 2. Few-shot with 5 corrected examples — This improved things, but now it overfits to the pattern and repeats the same CRUD structure even when I ask for a webhook handler. ## Does Retrieval-Augmented Generation Need Better Chunking? 3. RAG with the actual codebase — I indexed my project with langchain + chroma. Retrieval works, but the context window fills fast. The 7B model only has 4k context (8k if I push num_ctx), and the retrieved chunks eat 2k tokens before the prompt. ```python # Current workaround: post-generation lint loop import subprocess import ast ## Can Automated Validation Reliably Catch Invalid Code? def validate_python(code: str) -> tuple[bool, str]: try: ast.parse(code) result = subprocess.run( ["ruff", "check", "--select=F401,F821", "-"], input=code.encode(), capture_output=True, timeout=5 ) return result.returncode
Bigger context windows usually help, but they can fill up fast. For instance, when I used RAG with LangChain and ChromaDB, the retrieved chunks ate 2k tokens before the prompt, leaving little room for the model’s reasoning on a 7B setup. Which model are you running now?
Lowering the temperature usually helps with those fake endpoints. Have you tried a system prompt? You could try adding a system prompt with version-pinned docs by pasting the specific reference into the context window.
Few-shot prompts helped my 8B model avoid fake imports, though I still had to add a system prompt with version-pinned FastAPI 0.110 docs to catch deprecated exceptions like
RequestValidationError. I also wrote a post-generation lint loop withruffto catch missing imports (e.g.,Listfromtyping) before deployment, which saved me from runtime errors. Did that help your fake imports?