Develop Custom AI VS Code Extensions with LLM Integration

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Develop Custom AI VS Code Extensions with LLM Integration
Medium
~1-2 weeks
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1360
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1251
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    957
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Developers spend hours searching documentation and writing boilerplate code. With 5 years of experience and over 20 AI tools developed, we deliver robust VS Code extensions that embed directly into the editor and speed up routine tasks by 2–3 times. Integration with Anthropic and OpenAI APIs allows flexible model choice for each task, and streaming via client.messages.stream() delivers the first tokens in 200–300 ms. The average cost for one code explanation via Claude Haiku is about $0.25, and test generation via GPT-4o mini is $0.15 per 1K tokens. Typical extension development costs range from $3,000 to $15,000, and can save a development team 30-50% in code review time. For a team of 10 developers, our custom extension can save up to $5,000 annually.

One common issue is generation latency. We use streaming, with the first tokens appearing in 200–300 ms. P99 latency for short requests is under 1.5 s. To reduce hallucinations, we add context from the active file and few-shot examples — this decreases incorrect suggestions by 40%. We employ lexical analysis and AST traversal to extract context, and use tokenization strategies to fit within the model's context window. Inference optimization via model quantization further reduces latency. Our debounce mechanism and retry logic ensure robust performance.

What problems do we solve?

  • Generation latency — low-latency models (Claude Haiku, GPT-4o mini), streaming, and 300 ms debounce.
  • Hallucinations — project context, few-shot, custom prompts for your stack (Django, React, FastAPI).
  • Extension API integration — common mistakes: missing cleanup in deactivate(), WebView memory leaks. We address these during code review.
  • Marketplace publishing — preparing icons, description, testing on different VS Code versions (1–2 days).

How we do it

Stack: TypeScript, VS Code Extension API, Anthropic SDK / OpenAI SDK. For inline completion we use InlineCompletionItemProvider, for chat — WebView with acquireVsCodeApi(). All LLM calls are asynchronous, with retries on timeouts. Below is a comparison of models for different tasks.

Task Recommended Model P99 Latency Cost per 1K tokens
Explain code Claude Haiku 4 ~0.8 s ~$0.25
Refactor Claude Sonnet 4 ~1.2 s ~$3.0
Generate tests GPT-4o mini ~1.0 s ~$0.15
Chat history GPT-4o ~2.0 s ~$5.0
Approach Flexibility Performance Development Complexity
Built-in VS Code models Low High Low
LLM API (our implementation) High Medium (depends on model) Medium
Local LLM via ONNX Medium Low (depends on GPU) High

VS Code Extension API provides rich capabilities for integrating AI services. Documentation

Example OpenAI configuration
{
  "aiAssistant.model": "gpt-4o-mini",
  "aiAssistant.apiKey": "sk-..."
}

Why a custom extension is better than a ready-made one?

Ready-made solutions from the Marketplace often don't know the specifics of your stack. We customize prompts for your libraries and add code actions for specific errors. For example, the extension can automatically suggest try/except for blocks at risk of exceptions. This reduces code review time by 2–3 times.

What's included in the work?

  • Source code of the extension with documentation.
  • Integration with your LLM (we do not provide free API keys).
  • CI/CD setup for publishing.
  • Support for 1 month after delivery.

How the inline completion provider works

The provider implements the InlineCompletionItemProvider interface. On every character input, a request is sent to the LLM with the context of the current file. We use a 300 ms debounce and cache results for identical contexts. This keeps p99 latency below 1.5 s.

Extension structure

The code below is a configuration for four commands and code actions.

{
  "name": "ai-dev-assistant",
  "displayName": "AI Dev Assistant",
  "engines": { "vscode": "^1.85.0" },
  "activationEvents": ["onStartupFinished"],
  "contributes": {
    "commands": [
      { "command": "aiAssistant.explainCode", "title": "AI: Explain Code" },
      { "command": "aiAssistant.refactor", "title": "AI: Refactor Selection" },
      { "command": "aiAssistant.generateTests", "title": "AI: Generate Tests" },
      { "command": "aiAssistant.openChat", "title": "AI: Open Chat" }
    ],
    "keybindings": [
      { "command": "aiAssistant.explainCode", "key": "ctrl+shift+e", "when": "editorTextFocus" }
    ],
    "configuration": {
      "title": "AI Assistant",
      "properties": {
        "aiAssistant.apiKey": {
          "type": "string",
          "description": "Anthropic API Key"
        },
        "aiAssistant.model": {
          "type": "string",
          "default": "claude-haiku-4-5",
          "enum": ["claude-haiku-4-5", "claude-sonnet-4-5"]
        }
      }
    }
  },
  "main": "./out/extension.js"
}

Main extension file

import * as vscode from 'vscode';
import Anthropic from '@anthropic-ai/sdk';

let client: Anthropic;

export function activate(context: vscode.ExtensionContext) {
    const config = vscode.workspace.getConfiguration('aiAssistant');
    client = new Anthropic({ apiKey: config.get('apiKey') || '' });

    context.subscriptions.push(
        vscode.commands.registerCommand('aiAssistant.explainCode', explainSelectedCode)
    );
    // ... other commands
}

async function explainSelectedCode() {
    const editor = vscode.window.activeTextEditor;
    if (!editor) return;
    const selection = editor.selection;
    const selectedText = editor.document.getText(selection);
    if (!selectedText) {
        vscode.window.showWarningMessage('Select code to explain');
        return;
    }
    await vscode.window.withProgress(
        { location: vscode.ProgressLocation.Notification, title: 'AI analyzing code...' },
        async () => {
            const response = await client.messages.create({
                model: 'claude-haiku-4-5',
                max_tokens: 1024,
                messages: [{
                    role: 'user',
                    content: `Explain this code briefly and clearly:\n\`\`\`\n${selectedText}\n\`\`\``
                }]
            });
            const explanation = response.content[0].type === 'text' ? response.content[0].text : '';
            const outputChannel = vscode.window.createOutputChannel('AI Assistant');
            outputChannel.appendLine('=== AI Explanation ===');
            outputChannel.appendLine(explanation);
            outputChannel.show();
        }
    );
}

Chat Panel

class ChatPanel {
    private static currentPanel?: ChatPanel;
    private readonly panel: vscode.WebviewPanel;

    static createOrShow(extensionUri: vscode.Uri) {
        if (ChatPanel.currentPanel) {
            ChatPanel.currentPanel.panel.reveal();
            return;
        }
        const panel = vscode.window.createWebviewPanel(
            'aiChat', 'AI Chat', vscode.ViewColumn.Beside,
            { enableScripts: true }
        );
        ChatPanel.currentPanel = new ChatPanel(panel, extensionUri);
    }

    constructor(panel: vscode.WebviewPanel, extensionUri: vscode.Uri) {
        this.panel = panel;
        this.panel.webview.html = this.getWebviewContent();
        this.panel.webview.onDidReceiveMessage(async message => {
            if (message.type === 'chat') {
                const stream = await client.messages.stream({
                    model: 'claude-sonnet-4-5',
                    max_tokens: 2048,
                    messages: message.history,
                });
                for await (const chunk of stream.textStream) {
                    this.panel.webview.postMessage({ type: 'token', text: chunk });
                }
                this.panel.webview.postMessage({ type: 'done' });
            }
        });
    }

    private getWebviewContent(): string {
        return `<!DOCTYPE html>
<html>
<head><style>/* styles */</style></head>
<body>
    <div id="messages"></div>
    <input type="text" id="input" placeholder="Ask a question..." />
    <button onclick="sendMessage()">Send</button>
    <script>
        const vscode = acquireVsCodeApi();
        const history = [];
        function sendMessage() {
            const input = document.getElementById('input');
            history.push({ role: 'user', content: input.value });
            vscode.postMessage({ type: 'chat', history });
            input.value = '';
        }
        window.addEventListener('message', event => {
            const msg = event.data;
            if (msg.type === 'token') {
                // Append token
            }
        });
    </script>
</body>
</html>`;
    }
}

Code Actions Provider

class AICodeActionProvider implements vscode.CodeActionProvider {
    provideCodeActions(
        document: vscode.TextDocument,
        range: vscode.Range,
    ): vscode.CodeAction[] {
        const actions: vscode.CodeAction[] = [];
        const selectedText = document.getText(range);
        if (!selectedText) return actions;
        const explainAction = new vscode.CodeAction('AI: Explain', vscode.CodeActionKind.RefactorRewrite);
        explainAction.command = { command: 'aiAssistant.explainCode', title: 'Explain' };
        actions.push(explainAction);
        return actions;
    }
}

Work process

  1. Analysis — we study your scenarios, select the LLM model.
  2. Design — we design commands, WebView, providers.
  3. Development — we write TypeScript code, configure API clients.
  4. Testing — we test on real projects, measure latency.
  5. Publishing — we publish to the VS Code Marketplace, set up CI/CD.

Approximate timelines

  • Basic commands (explain, refactor): 3–5 days.
  • Chat panel with WebView: 1 week.
  • Inline completion provider: 1–2 weeks.
  • Marketplace publishing: 1–2 days.

Typical mistakes in self-development

  • Using activate() without calling context.subscriptions.push() — commands are not registered.
  • Ignoring dispose() for WebView — memory leak.
  • No fallback when API is unavailable — user sees a blank screen.

Contact us to evaluate your project. Order the development of an AI extension that will work specifically for your tasks. Our team's experience: 5 years, we guarantee quality and support for 30 days after publication. Tell us about your project — we'll suggest the optimal extension architecture. For an Anthropic VS Code extension, we configure the SDK accordingly; similarly for a Claude VS Code extension. The Inline Completion VS Code feature is implemented via InlineCompletionItemProvider. Our AI code autocomplete uses low-latency models. This AI-powered code extension is fully customizable.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.