Industrial Prompt Engineering: Stable LLM Responses Turnkey

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Industrial Prompt Engineering: Stable LLM Responses Turnkey
Simple
from 1 day to 3 days
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Implementing Prompt Engineering for an AI System

Deployed an LLM in a support chatbot and got 40% hallucinations? A familiar situation. One of our clients, a fintech startup, spent a month developing a chatbot for credit consultations. A junior wrote the prompt in an hour—result: 40% of responses contained incorrect data on interest rates and terms. In two weeks, we designed a multi-level prompt with anti-hallucination instructions, chain-of-verification, and few-shot examples. Hallucinations dropped to 3%, accuracy increased from 55% to 94%, and p95 latency halved. In 5 years of working with AI systems, we have developed an approach that delivers predictable p95 latency and >95% accuracy under production loads.

Prompt Engineering is the discipline of designing input data for LLMs to obtain predictable, high-quality results. It includes structuring queries, managing context, selecting techniques (CoT, Few-Shot, ReAct), tuning parameters, and iterative calibration.

Why Standard Prompts Fail in Production

A common mistake: prompts written by a developer in ten minutes. They give an acceptable answer on 2–3 manual tests but fail under real load. Main issues:

  • Hallucination—the model makes up things not in context. Without anti-hallucination instructions, up to 30% of responses contain false facts.
  • Context sensitivity—a small wording change completely alters the response. Temperature and top-p not tied to the task type create irreproducibility.
  • Lack of error handling—when data is missing, the model does not say "I don't know" but invents an answer.

We eliminate these problems through structured templates, verification, and A/B tests.

How to Reduce the Hallucination Percentage in LLMs?

Key technique: chain-of-verification—the model first generates an answer, then checks each claim against the context. We add a system prompt that prohibits inventing and introduces a confidence threshold. Additionally, we use few-shot examples with edge cases. In production, this reduces the hallucination rate from 20% to 2–3%. According to OpenAI recommendations, structured prompts cut hallucinations by 30–50%.

ANTI_HALLUCINATION_ADDENDUM = """
IMPORTANT: Only answer based on the provided context.
If the information is not in the context, say: "I have no data on this question."
Do not guess or make assumptions.
If confidence is < 80%, indicate the degree of uncertainty.
"""

async def answer_with_verification(question: str, context: str) -> dict:
    answer = query_llm(
        f"Context:\n{context}\n\nQuestion: {question}",
        system=f"You are an analyst. {ANTI_HALLUCINATION_ADDENDUM}",
    )
    verification = query_llm(
        f"Original answer: {answer}\n\nQuestion: Are all statements in the answer confirmed by the context? Answer as JSON: {{\"verified\": bool, \"unsupported_claims\": [...]}}",
        temperature=0,
    )
    return {"answer": answer, "verification": json.loads(verification)}

How We Design Industrial Prompts

The process starts not with code but with analyzing expectations: what the prompt should do, boundary cases, how to measure quality. We collect 100–200 representative queries and label the ideal answer. Only then do we write the first baseline.

Role Model and Context

The system prompt is built using the scheme: "You are [role]. Task: [goal]. Context: [conditions, data]. Rules: [what is allowed, what is not]. Output format: [explicit format]." For RAG, we add a source pointer. Example:

SYSTEM_PROMPT = """You are {role}.

Task: {task_description}

Rules:
{rules}

Output format:
{output_format}"""

A/B Testing and Metrics

We test each prompt variant on an eval set using LLM-as-judge. Compare by accuracy, completeness, style. Example test class:

class PromptABTest:
    def __init__(self, variants: dict[str, str]):
        self.variants = variants
        self.results = {name: [] for name in variants}

    def run_test(self, test_inputs: list[str], judge_prompt: str) -> dict:
        for input_text in test_inputs:
            outputs = {}
            for name, prompt in self.variants.items():
                output = query_llm(input_text, system=prompt)
                outputs[name] = output

            comparison = query_llm(
                f"""Compare two responses to the query: \"{input_text}\"

Variant A: {outputs[list(outputs.keys())[0]]}
Variant B: {outputs[list(outputs.keys())[1]]}

{judge_prompt}
Return JSON: {{\"winner\": \"A\"|\"B\"|\"tie\", \"reason\": \"...\"}}""",
                temperature=0,
            )
            result = json.loads(comparison)
            winner = result["winner"]
            if winner != "tie":
                winning_name = list(self.variants.keys())[0 if winner == "A" else 1]
                self.results[winning_name].append(1)

        return {name: sum(wins) / len(test_inputs) for name, wins in self.results.items()}

Step-by-Step Prompt Calibration Guide

  1. Define the goal and metrics: accuracy, completeness, latency, hallucination rate.
  2. Collect an eval set: 100–200 real queries with reference answers.
  3. Create a baseline: a simple prompt with role model and basic rules.
  4. Iteratively improve: for each problem, add instructions, few-shot examples, checks.
  5. A/B testing: compare variant vs control on shadow traffic, pick the best.

Comparison of Prompt Engineering Techniques

Technique Application Effect
Few-Shot Add 3–5 examples to the prompt Improves accuracy by 15–25%
Chain-of-Thought Step-by-step reasoning Boosts complex task quality by 30%
Self-Consistency Multiple generations + voting Reduces variance, increases reliability
Chain-of-Verification Fact-checking after the answer Cuts hallucination rate by 5–10 times

What's Included in the Work

After calibration, you receive:

  • Prompt templates in JSON/YAML with comments.
  • An eval set of 200+ examples with reference answers.
  • A metrics report: accuracy, recall, hallucination rate, p99 latency.
  • A/B testing on shadow traffic (optional).
  • Team training: a 2–4 hour workshop on maintaining and refining prompts.
  • Stability guarantee: if quality drops after deployment, we adjust the prompt free of charge within a month.

Approach Comparison: One-shot vs Iterative

Characteristic One-shot (manual) Iterative (with eval) With A/B testing
Time to production 1 day 3–5 days 5–10 days
Accuracy on test set 30–50% 60–80% 90–95%
Latency stability Low (p99 grows) Medium High (fixed p99)
Hallucination risk High (>15%) Medium (5–10%) Low (<3%)

An iterative approach with metrics is 2–3 times more effective than one-off writing—confirmed on 50+ projects.

Timeline and Results

  • Basic prompt for a specific use case: 1–3 days.
  • A/B testing with an eval set: 3–5 days.
  • Production prompt with verification and documentation: 7–10 business days.

Cost is calculated individually, depending on the task complexity and number of prompts. A typical project ranges from $1500 to $4000. Savings from fixing hallucinations can reach $10,000 per month. Estimation takes 1 day: you describe the use case, we analyze and propose a plan.

Want stable LLM operation? Contact us to analyze your current prompt and suggest improvements within 1 day. Order a prompt audit and get a detailed report with metrics and recommendations. Get a consultation on your prompt today.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.