Your Django team of 12 developers spends hours on repetitive code reviews. We built a custom AI assistant that understands your codebase and delivers relevant suggestions with <200ms latency. Here's how.
A custom AI assistant for the IDE is not just autocomplete on steroids. It keeps the entire project context: open files, change history, database schema, tests. A properly built assistant knows you're writing a user registration function in a Django project with PostgreSQL and suggests code compatible with your models and conventions.
Problems we solve
Generic models don't know project context. GitHub Copilot gives average-quality suggestions, ignoring internal APIs, custom ORM methods, and architectural decisions. The acceptance rate of such suggestions rarely exceeds 23%.
Code confidentiality. Teams with NDAs cannot send code to cloud services. A fully local stack is required.
Suggestion latency. Cloud solutions often have latency >500ms, killing the magic. For inline completion, latency <200ms is critical.
A custom assistant solves all three: it uses your codebase, works locally, and delivers suggestions in 80–150ms.
Architecture of an IDE assistant
A full Copilot-like assistant consists of several layers:
- Context Collector — gathers relevant context: current file, imports, related files, cursor position, selected code, clipboard.
- LSP Bridge — interacts with the Language Server Protocol to get AST, types, definitions.
- Retrieval Engine — semantic search over the codebase using embeddings (CodeBERT, text-embedding-3-small) and a vector store with RAG.
- LLM Gateway — request routing: fast model for inline completion, powerful model for chat/refactoring.
- Response Renderer — output formatting: diff for refactoring, ghost text for completion, markdown for chat.
Why a custom AI assistant outperforms GitHub Copilot
A custom assistant uses your project's context: codebase indexes, DB schemas, issue trackers. This yields more relevant suggestions than generic models. In our case study, acceptance rate rose from 23% to 41%, and subscription costs were cut in half (saving $2,000 per month). Plus, you have full data control — no code leaks to cloud services.
Continue.dev — open-source foundation
Continue.dev (https://github.com/continuedev/continue) is the most mature open-source alternative to GitHub Copilot. It supports VS Code and JetBrains, configurable via ~/.continue/config.json.
{
"models": [
{
"title": "Claude 3.5 Sonnet",
"provider": "anthropic",
"model": "claude-sonnet-4-5",
"apiKey": "$ANTHROPIC_API_KEY"
},
{
"title": "Ollama Qwen2.5-Coder",
"provider": "ollama",
"model": "qwen2.5-coder:7b",
"apiBase": "http://localhost:11434"
}
],
"tabAutocompleteModel": {
"title": "Autocomplete",
"provider": "ollama",
"model": "qwen2.5-coder:1.5b"
},
"contextProviders": [
{"name": "code", "params": {}},
{"name": "docs", "params": {}},
{"name": "diff", "params": {}},
{"name": "terminal", "params": {}},
{"name": "problems", "params": {}},
{"name": "folder", "params": {}},
{"name": "codebase", "params": {}}
],
"slashCommands": [
{"name": "edit", "description": "Edit highlighted code"},
{"name": "comment", "description": "Write comments for the code"},
{"name": "tests", "description": "Write unit tests"},
{"name": "share", "description": "Export the chat session"}
]
}
Key feature: tabAutocompleteModel uses a fast local model (1.5B parameters), while chat uses a powerful cloud model. Inline completion latency: 80–150ms on Qwen2.5-Coder 1.5B via Ollama.
Custom context provider: example for database schema
Continue.dev allows writing custom context providers for specific data sources:
import { ContinueConfig, IContextProvider } from "@continuedev/core";
class DatabaseSchemaProvider implements IContextProvider {
get description() {
return { title: "db", displayTitle: "Database Schema", description: "Current database schema", type: "normal" };
}
async getContextItems(query: string, extras: any) {
const schema = await fetchDatabaseSchema();
return [{ name: "Database Schema", description: "Current DB schema", content: schema }];
}
}
export function modifyConfig(config: ContinueConfig): ContinueConfig {
config.contextProviders = [...(config.contextProviders || []), new DatabaseSchemaProvider()];
return config;
}
This allows the assistant to consider table structures, foreign keys, and indexes when generating queries.
How we configure context-aware suggestions for your project
The setup process consists of four steps.
-
Codebase analysis: we identify key patterns, internal APIs, and database structure. We use a static analyzer to extract metadata.
-
Custom context providers: for each source (DB schema, Jira, documentation) we write a provider in TypeScript or Python. An example for DB schema is shown above.
-
Indexing with RAG: we build a semantic index of the code using embeddings (CodeBERT or text-embedding-3-small) and a vector database (ChromaDB, pgvector). The index updates on repository pushes.
-
Fine-tuning (optional): we fine-tune the model on your historical PRs and typical tasks to improve suggestion relevance. We use LoRA to save resources.
As a result, the assistant suggests code that follows your conventions, not abstract examples.
Practical case: rollout to a 12-developer team
Starting state: team used GitHub Copilot, complained about irrelevant suggestions — Copilot didn't know internal patterns of a Django project with 800+ models.
Solution: Continue.dev + local Ollama for autocomplete + Claude via API for chat/refactoring + custom context provider with codebase index.
Infrastructure: server with RTX 4090 (Qwen2.5-Coder 7B for autocomplete), Claude API for complex requests.
Results after 2 months:
- Inline suggestion acceptance: 23% (Copilot) → 41% (custom) — 1.78x better than Copilot.
- Average time to write a typical CRUD endpoint: 52 min → 31 min (40% faster).
- Tasks like "write a test for this function": 100% manual → 70% automated.
- Subscription savings: over 50%, saving $2,500 per month.
Key factor for acceptance rate improvement: the context provider with codebase index gave the model real examples from the project, not abstract code.
Local models for completion
For teams with code confidentiality requirements — a fully local stack. To maximize GPU utilization we use INT4 quantization.
| Model | Size | Latency (RTX 3080) | Quality |
|---|---|---|---|
| Qwen2.5-Coder 1.5B | 1.5B | 50–80 ms | Basic |
| Qwen2.5-Coder 7B | 7B | 150–250 ms | Good |
| DeepSeek-Coder 6.7B | 6.7B | 140–230 ms | Good |
| CodeLlama 13B | 13B | 350–500 ms | High |
For inline completion, latency <200 ms is critical — users notice delay. Therefore models up to 7B are used for FIM (fill-in-the-middle).
Timelines and process
| Stage | What we do | Duration |
|---|---|---|
| Analysis | Audit codebase, identify key patterns | 1–2 days |
| Configuration | Set up Continue.dev, select and connect models | 2–3 days |
| Development | Custom context providers (DB, Jira, docs) | 1 week |
| Indexing | Semantic index + code vectorization | 1–2 weeks |
| Onboarding | Team training, configure rules and templates | 1 week |
| Support | Warranty and technical support for one month | — |
What's included
- Configuration and architecture documentation
- Access to selected models (local or cloud)
- Team training (2-hour workshop)
- Technical support and one-month warranty
- Source code for custom context providers (if developed)
Total: 3–5 weeks to full implementation. Pricing starts from $15,000 for a standard team. Contact us to get a consultation and project estimate.
We help teams of any size, from startups to enterprise with custom security requirements. We have over 5 years of experience in AI/ML and 20 implemented projects. We provide a warranty on integration and post-implementation support.
Order a custom AI assistant for your IDE — reach out, and we'll tell you in detail how to accelerate your development. Get a consultation — we'll evaluate your project.







