We often encounter a situation where a team of 4–5 developers spends 4–6 hours creating a single standard CRUD endpoint with tests. Meanwhile, the business demands 3–5 new endpoints per week — routine takes 70% of the time, leaving little for business logic and architecture. This is the problem solved by an AI code generation system: it takes over the boilerplate, leaving the engineer with creative tasks.
How AI Code Generation Accelerates Development
A custom AI agent doesn't just insert code from a template — it understands the context of your codebase: DB schemas, existing classes, API contracts, code style. Based on that, it generates production-quality code, checks syntax, runs tests, and iteratively fixes errors. Architecturally, such a system includes:
- Context Manager — collects relevant context: DB schema, API interfaces, existing models, code style guide.
- Generation Engine — LLM agent with tools for reading files, running tests, searching the codebase.
- Verification Layer — syntax checking, test execution, linter.
- Feedback Loop — iterations based on test errors.
Code Generation Agent on LangGraph
from langgraph.graph import StateGraph, END
from langchain_openai import ChatOpenAI
from langchain_core.tools import tool
from typing import TypedDict, Annotated, Optional
import subprocess
import ast
import operator
llm = ChatOpenAI(model="claude-opus-4-5", temperature=0.1)
class CodeGenState(TypedDict):
task_description: str
existing_code_context: str
generated_code: Optional[str]
test_results: Annotated[list, operator.add]
iteration: int
max_iterations: int
errors: Annotated[list, operator.add]
final_code: Optional[str]
@tool
def read_file(file_path: str) -> str:
"""Read a file from the codebase to get context."""
try:
with open(file_path) as f:
return f.read()
except FileNotFoundError:
return f"File {file_path} not found"
@tool
def search_codebase(query: str, directory: str = "./src") -> str:
"""Search the codebase with grep to find similar code."""
result = subprocess.run(
["grep", "-r", "--include=*.py", "-n", query, directory],
capture_output=True, text=True
)
return result.stdout[:3000] if result.stdout else "Nothing found"
@tool
def run_python_syntax_check(code: str) -> str:
"""Check Python code syntax."""
try:
ast.parse(code)
return "Syntax is correct"
except SyntaxError as e:
return f"Syntax error: {e}"
@tool
def run_tests(test_file_path: str) -> str:
"""Run pytest and return results."""
result = subprocess.run(
["python", "-m", "pytest", test_file_path, "-v", "--tb=short"],
capture_output=True, text=True, timeout=60
)
output = result.stdout + result.stderr
return output[-3000:] # Last 3000 characters
@tool
def write_file(file_path: str, content: str) -> str:
"""Write code to a file."""
with open(file_path, "w", encoding="utf-8") as f:
f.write(content)
return f"File {file_path} written ({len(content)} characters)"
CODE_GEN_SYSTEM = """You are a Senior Software Engineer. Generate production-quality code.
Principles:
- Follow existing codebase patterns
- Write typed code (type hints)
- Each function has one level of abstraction
- Handle errors explicitly
- Minimize dependencies on external libraries if standard alternatives exist
Process:
1. Read existing code for context
2. Generate code in the same style
3. Check syntax
4. Run tests
5. Fix errors iteratively"""
from langgraph.prebuilt import create_react_agent
code_gen_agent = create_react_agent(
llm.bind_tools([read_file, search_codebase, run_python_syntax_check, run_tests, write_file]),
tools=[read_file, search_codebase, run_python_syntax_check, run_tests, write_file],
state_modifier=CODE_GEN_SYSTEM,
)
Why Codebase Context Matters
Without context, LLMs generate code that doesn't fit into the existing architecture — different style, wrong names, incompatible imports. Our Context Aware Code Generator automatically collects relevant files: data models, base classes, code style guide. This is critical for projects using FastAPI + SQLAlchemy with custom patterns. Example implementation:
class ContextAwareCodeGenerator:
def __init__(self, project_root: str):
self.project_root = project_root
self.context_cache = {}
async def gather_context(self, task: str) -> str:
"""Gather relevant context for the task"""
# Find similar files via LLM
relevant_files = await self.identify_relevant_files(task)
context_parts = []
# Read DB schema
if await self.file_exists("models.py"):
models = await read_file_async(f"{self.project_root}/models.py")
context_parts.append(f"## Data Models\n{models[:2000]}")
# Read base classes and interfaces
for file_path in relevant_files[:3]:
content = await read_file_async(file_path)
context_parts.append(f"## {file_path}\n{content[:1500]}")
# Add code style guide
if await self.file_exists(".codestyle.md"):
style = await read_file_async(f"{self.project_root}/.codestyle.md")
context_parts.append(f"## Code Style\n{style[:1000]}")
return "\n\n".join(context_parts)
async def generate(self, task: str, output_file: str) -> dict:
context = await self.gather_context(task)
result = await code_gen_agent.ainvoke({
"messages": [{
"role": "user",
"content": f"""Task: {task}
Codebase context:
{context}
Output file: {output_file}
Generate the code, check it, and write to file."""
}]
})
return {
"task": task,
"output_file": output_file,
"iterations": result.get("iteration", 1),
"tests_passed": self.extract_test_status(result),
}
Template-based Generation with LLM Filling
For typical tasks (CRUD, migrations, tests), a hybrid approach is effective: a template with placeholders that the LLM expands. This provides predictable structure and control over critical parts.
class CRUDGenerator:
"""Generates CRUD modules from entity schema"""
CRUD_TEMPLATE = """
# Module for entity {entity_name}
from sqlalchemy import Column, Integer, String, DateTime, func
from sqlalchemy.orm import Session
from pydantic import BaseModel
from typing import Optional, List
from datetime import datetime
# PLACEHOLDERS FOR LLM REPLACEMENT:
# COLUMNS - list of SQLAlchemy columns
# PYDANTIC_FIELDS - Pydantic schema fields
# BUSINESS_LOGIC - specific business logic
"""
async def generate_crud_module(self, entity_spec: dict) -> str:
"""entity_spec: {name, fields, business_rules, relationships}"""
# LLM fills specific parts
columns = await self.generate_sqlalchemy_columns(entity_spec["fields"])
schemas = await self.generate_pydantic_schemas(entity_spec["fields"])
business_logic = await self.generate_business_logic(entity_spec.get("business_rules", []))
# Assemble final module
result = await llm.ainvoke(f"""Create a full CRUD module for entity {entity_spec['name']}.
Specification: {json.dumps(entity_spec, ensure_ascii=False)}
Stack: FastAPI + SQLAlchemy 2.0 + Pydantic v2
Include: model, pydantic schemas, CRUD functions, FastAPI router with dependency injection
Code standards: async/await, type hints, docstrings""")
return result.content
Case Study: Fintech Startup
Our client — a fintech company with 4 developers — spent 4–6 hours on a standard CRUD endpoint with tests. After implementing AI generation, the time dropped to 50 minutes (15 min generation + 35 min review). Metrics before and after:
| Metric | Before AI | After AI | Improvement |
|---|---|---|---|
| Time per CRUD endpoint | 5 hours | 50 minutes | 6x faster |
| Test coverage of new endpoints | 45% | 82% | +37 pp |
| Code consistency | Low (different patterns) | High (single pattern) | Significant |
| Post-generation rework needed | — | 14% of PRs | 86% accepted without changes |
The situation allowed the team to save significant costs on routine tasks, and the cost per PR decreased several times. The system generated CRUD modules from OpenAPI specifications, automatically created pytest tests and Alembic migrations. An AI Code Review agent provided suggestions at the review stage. The only challenge was business logic in 14% of cases requiring substantial rework, which is addressed by adding rules to the context.
Model Comparison for Code Generation
| Model | Latency (p99) | Code Quality | Cost |
|---|---|---|---|
| GPT-4o | 2.1 s | Excellent | Medium |
| Claude 3.5 Sonnet | 3.8 s | Outstanding | High |
| LLaMA 3 (70B, INT4) | 0.9 s | Good | Low |
| Mistral (7B, INT4) | 0.4 s | Average | Very Low |
Model selection depends on budget and quality requirements. For production tasks, we recommend Claude 3.5 Sonnet or GPT-4o with large context window.
What's Included
When ordering our service, you receive:
- Codebase and architecture audit — assessment of CI/CD maturity, code style, test coverage.
- Design and implementation of an AI agent — tailored to your stack and requirements.
- Integration with IDE (VS Code, JetBrains) and CI/CD (GitHub Actions, GitLab CI, Jenkins).
- Team training — workshops on using the system, best practices for prompting.
- Documentation and 1 month warranty support.
Order AI generation implementation and accelerate your development.
Estimated Timelines
- Basic generator with context: 2–3 weeks
- Agentic loop with tests and iterations: 2–3 weeks
- Integration into CI/CD and IDE: 2–3 weeks
- Total: 6–9 weeks
Pricing is calculated individually, based on scope and complexity. Contact us for a project assessment — we'll propose the optimal solution. Our LangGraph agent is already running in several production systems, we guarantee stability and support.







