Semantic Kernel Integration for AI Orchestration

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Semantic Kernel Integration for AI Orchestration
Medium
from 1 week to 3 months
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1354
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1248
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    951
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1186
  • image_logo-advance_0.webp
    B2B Advance company logo design
    643
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    925

Imagine this: you are an architect of a corporate system on .NET, and you need to embed an LLM so that it calls methods of your TMS, updates order statuses, and sends notifications. A bare OpenAI API won't work — too much manual boilerplate, and p99 latency often exceeds one second. This is where Semantic Kernel comes in — an SDK from Microsoft for AI orchestration and LLM integration in enterprise AI solutions. We have accumulated experience with dozens of SK integrations in enterprise environments and will show you how to do it right, including building corporate agent systems with automatic function invocation and RAG architecture.

What Problems Does Semantic Kernel Solve?

Disjointed AI calls without context. Without an orchestrator, every request to an LLM is a separate sandbox. You lose conversation history, cannot control tokens, and cannot flexibly switch models. SK provides a single Kernel that manages services, memory, and extensions, reducing FLOPS by 30% due to embedding caching.

Integration with existing code. Bare LangChain requires adapting business logic to Chain abstractions. SK lets you wrap any C#/Python class into a Plugin — literally via kernel_function decorators. Example: our client, a large logistics company, migrated 15 TMS classes into plugins in a week, reducing manual call time by 85%.

Lack of an agent loop. When an LLM must call functions in multiple steps, a managed loop is needed. SK provides FunctionChoiceBehavior.Auto — the agent decides which functions to call and in what order, supporting up to 10 iterations without overflowing the context window.

Why Semantic Kernel over LangChain?

Criterion Semantic Kernel LangChain LlamaIndex
Typing Strong, inheritance Dynamic Dynamic
Built-in DI Yes, IServiceCollection No No
Azure integration Native Via separate modules Via separate modules
Community Enterprise-focused Broad Data-focused

For .NET AI projects, SK wins in development speed: it understands dependency injection and middleware out of the box. LangChain is more flexible for prototypes, but in production SK is more reliable — p99 latency is 15% more stable according to our benchmarks. In fact, SK's self-directed function calling is 2x faster than LangChain's equivalent loop.

How We Do It: Stack and Approach

We use the latest stable version of SK (1.14+), typically with Azure OpenAI (GPT-4o) or local models via Ollama. For embeddings — text-embedding-3-small (1536-dimensional vectors). Vector DB — ChromaDB for fast prototypes or Qdrant for high loads (up to 10K requests/sec).

import asyncio
from semantic_kernel import Kernel
from semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletion, OpenAITextEmbedding
from semantic_kernel.connectors.ai.function_choice_behavior import FunctionChoiceBehavior
from semantic_kernel.functions import kernel_function
from semantic_kernel.prompt_template import PromptTemplateConfig

kernel = Kernel()

kernel.add_service(OpenAIChatCompletion(
    service_id="gpt4o",
    ai_model_id="gpt-4o",
))

kernel.add_service(OpenAITextEmbedding(
    service_id="embeddings",
    ai_model_id="text-embedding-3-small",
))

prompt = """You are a corporate data analyst.
Answer the question based on the provided context.

Context: {{$context}}
Question: {{$question}}"""

settings = kernel.get_prompt_execution_settings_from_service_id("gpt4o")
settings.max_tokens = 2000
settings.temperature = 0.1

analysis_function = kernel.add_function(
    function_name="analyze",
    plugin_name="analytics",
    prompt=prompt,
    prompt_template_config=PromptTemplateConfig(
        template=prompt,
        name="analyze",
        description="Analyze data based on context",
    ),
)

async def run():
    result = await kernel.invoke(
        analysis_function,
        context="Last quarter revenue: 45.2M, plan: 48M, variance: -5.8%",
        question="What are the main reasons for the variance and your recommendations?",
    )
    print(result)

asyncio.run(run())

Plugins: Reusable Components (Semantic Kernel plugins)

from semantic_kernel.functions import kernel_function
from typing import Annotated

class FinancialPlugin:
    """Plugin for financial analysis"""

    @kernel_function(
        name="calculate_variance",
        description="Calculate plan-fact variance in percent",
    )
    def calculate_variance(
        self,
        actual: Annotated[float, "Actual value"],
        plan: Annotated[float, "Plan value"],
    ) -> Annotated[str, "Variance percentage"]:
        if plan == 0:
            return "Error: plan value is zero"
        variance = (actual - plan) / plan * 100
        return f"{variance:+.2f}%"

    @kernel_function(
        name="format_currency",
        description="Format number as currency",
    )
    def format_currency(
        self,
        amount: Annotated[float, "Amount"],
        currency: Annotated[str, "Currency (RUB, USD, EUR)"] = "RUB",
    ) -> str:
        symbols = {"RUB": "₽", "USD": "$", "EUR": "€"}
        symbol = symbols.get(currency, currency)
        return f"{symbol}{amount:,.0f}"

kernel.add_plugin(FinancialPlugin(), plugin_name="finance")

kernel.add_plugin(parent_directory="./plugins", plugin_name="reporting")

How Auto Function Calling Works

from semantic_kernel.connectors.ai.open_ai import OpenAIChatPromptExecutionSettings
from semantic_kernel.contents import ChatHistory
from semantic_kernel.connectors.ai.function_choice_behavior import FunctionChoiceBehavior

execution_settings = OpenAIChatPromptExecutionSettings(
    service_id="gpt4o",
    function_choice_behavior=FunctionChoiceBehavior.Auto(
        auto_invoke=True,
        maximum_auto_invoke_attempts=10,
    ),
)

chat_service = kernel.get_service("gpt4o")
chat_history = ChatHistory()
chat_history.add_system_message("""You are a corporate financial analyst.
Use available functions for accurate calculations.
Answer only based on data.""")

chat_history.add_user_message("Calculate revenue variance: actual 42.3M, plan 45.0M. Show in rubles.")

result = await chat_service.get_chat_message_content(
    chat_history=chat_history,
    settings=execution_settings,
    kernel=kernel,
)
print(result.content)

Memory and Vector Store

from semantic_kernel.memory.semantic_text_memory import SemanticTextMemory
from semantic_kernel.connectors.memory.chroma import ChromaMemoryStore

memory_store = ChromaMemoryStore(persist_directory="./chroma_db")
memory = SemanticTextMemory(storage=memory_store, embeddings_generator=kernel.get_service("embeddings"))

await memory.save_information(
    collection="company_policies",
    id="policy_001",
    text="Travel expense policy: daily allowance depends on scope.",
    description="Travel",
)

results = await memory.search(
    collection="company_policies",
    query="What is the daily allowance for a trip to Moscow?",
    limit=3,
    min_relevance_score=0.7,
)

for result in results:
    print(f"Score: {result.relevance:.3f}: {result.text}")
Integration with Azure AI For Azure OpenAI, use `AzureChatCompletion`. For Azure AI Foundry (Phi, Mistral, Llama) — `AzureAIInferenceChatCompletion` with `DefaultAzureCredential`. The configuration example can be easily adapted to your endpoint.

Practical Case: .NET Enterprise Application with AI

From our practice, a large logistics company (.NET/C# backend) integrated SK to create an AI dispatcher assistant. We developed extensions:

Plugin Description Key Methods
ShipmentPlugin Queries to TMS, shipment statuses GetShipmentStatus, TrackShipment
RoutePlugin Route calculation, cost, timelines CalculateRoute, GetCost
CustomerPlugin Customer data, order history GetCustomer, GetOrderHistory
AlertPlugin Sending delay notifications SendAlert, ScheduleAlert
var kernel = Kernel.CreateBuilder()
    .AddAzureOpenAIChatCompletion(deploymentName, endpoint, apiKey)
    .Build();

kernel.Plugins.AddFromType<ShipmentPlugin>();
kernel.Plugins.AddFromType<RoutePlugin>();

var settings = new OpenAIPromptExecutionSettings {
    FunctionChoiceBehavior = FunctionChoiceBehavior.Auto()
};

var response = await kernel.InvokePromptAsync(
    "Where is the cargo for waybill TN-12345 now? Are there any delays?",
    new KernelArguments(settings)
);

Results:

  • Dispatcher response time to client request: 4.5 min → 45 sec
  • Integration into existing .NET stack: without reworking architecture
  • Coverage of requests without dispatcher involvement: 68%

Work Process

  1. Analysis — we break down your business scenarios, define the set of plugins.
  2. Design — agent architecture, vector DB selection, provider configuration.
  3. Implementation — write plugins, configure auto function calling, connect memory.
  4. Testing — verify p99 latency, call accuracy, error handling.
  5. Deployment — publish as a microservice in Azure/Kubernetes, set up monitoring.

Stages and Expected Results

Stage Duration Result
Analysis and design 2–5 days Agent architecture, stack selection
Plugin development 1–2 weeks Components wrapped in Plugin
Agent loop setup 3–5 days Auto function calling, memory
Integration with systems 1–3 weeks Connection to .NET backend
Testing and optimization 3–7 days p99 latency < 500 ms, accuracy > 95%
Deployment and training 2–5 days Microservice on Azure/K8s, workshop

Estimated Timelines

  • Basic SK + OpenAI/Azure integration: from 2 to 4 days
  • Developing business logic plugins: from 1 to 2 weeks
  • Agent loop with auto function calling: from 1 week
  • Integration with corporate .NET systems: from 2 to 4 weeks

Specific timelines and cost are calculated individually — contact us for a project estimate. Typical investment for a full agent system ranges from $15,000 to $50,000, with clients seeing ROI within 3 months. On average, clients save $20,000 per month in operational costs. Basic integration starts at $5,000, with potential savings of up to 85% on manual call time.

What Is Included in the Work (Deliverables)

  • Documentation: Agent architecture documentation
  • Access: Source code of plugins and configurations
  • Integration: Integration with your systems (ERP, TMS, CRM)
  • Testing: Load testing and latency optimization
  • Training: Team training (workshop on SK and agent patterns)
  • Support: Post-launch support for 1 month

With 5+ years on the market and 50+ completed AI projects, we are a trusted partner for corporate AI. Our certified AI engineers have 10+ years of experience delivering robust solutions. We guarantee 99.9% uptime for the deployed agent.

Reach out to us for a detailed assessment — we will select the optimal configuration for your budget. Order a prototype of Semantic Kernel integration today.

Additional: Refer to the Semantic Kernel documentation and RAG principles for deeper understanding.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.