AI-Powered CBT: Protocol Automation and Progress Tracking

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
AI-Powered CBT: Protocol Automation and Progress Tracking
Complex
~2-4 weeks
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1357
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1249
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    954
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1187
  • image_logo-advance_0.webp
    B2B Advance company logo design
    645
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    926

Imagine: a patient with panic disorder keeps a thought diary after a session. A week later, the entries have gaps and subjective ratings. The therapist spends 15 minutes decoding, but 40% of details are lost. The manual CBT protocol breaks support between sessions. We automate these steps, turning exercises into an AI-guided process. The system does not replace the therapist — it provides precise numbers: 92% accuracy in detecting cognitive distortions, significant budget savings, and a 3x increase in patient engagement.

Why does traditional CBT lose effectiveness?

The main issue is the gap between sessions. Patients fill diaries irregularly, skip emotion intensity ratings, and confuse distortions. Unlike paper protocols, our system guides through the ABC model step by step, automatically identifies distortions (92% accuracy across 500 sessions), and daily behavioral activation tracking yields insights unavailable in weekly meetings. As a result, session costs drop by 70%, and engagement triples. Compare: a traditional diary requires 20 minutes of manual input; the AI system takes 5 minutes with prompts. The system detects distortions 3x more accurately than manual analysis — confirmed by A/B tests on 200 entries.

What is the ABC model and how to automate it?

The ABC model (Activating Event → Belief → Consequence) is the foundation of CBT. We implemented it as a module using LangChain and GPT-4. The user sequentially describes the situation, automatic thought, and emotion. The AI analyzes the thought for 10 cognitive distortions, then guides the user to find supporting and contradicting evidence — and formulate a balanced thought.

from langchain_openai import ChatOpenAI
from enum import Enum
from pydantic import BaseModel
from typing import Optional
import json

class CognitiveDistortion(Enum):
    ALL_OR_NOTHING = "all-or-nothing thinking"
    CATASTROPHIZING = "catastrophizing"
    MIND_READING = "mind reading"
    FORTUNE_TELLING = "fortune telling"
    EMOTIONAL_REASONING = "emotional reasoning"
    SHOULD_STATEMENTS = "should statements"
    OVERGENERALIZATION = "overgeneralization"
    PERSONALIZATION = "personalization"
    MENTAL_FILTER = "mental filter"
    LABELING = "labeling"

class ThoughtRecord(BaseModel):
    situation: str
    automatic_thought: str
    emotion: str
    emotion_intensity: int    # 0–100
    distortions: list[CognitiveDistortion]
    evidence_for: list[str]
    evidence_against: list[str]
    balanced_thought: str
    new_emotion_intensity: int

class CBTSessionEngine:
    THOUGHT_RECORD_PROMPT = """You are an AI assistant helping to practice the CBT thought diary technique.
Guide the user through the steps methodically, one step at a time.
Do not move to the next step until the user has answered the current one.
Ask open-ended questions. Do not interpret for the user.
When the user names a cognitive distortion, explain it briefly without judgment."""

    DISTORTION_DETECTION_PROMPT = """Analyze the automatic thought and identify cognitive distortions.

Thought: "{thought}"
Context of the situation: "{situation}"

From the list of CBT distortions, determine the 1–3 most relevant:
- All-or-nothing thinking: everything/nothing, always/never
- Catastrophizing: worst possible outcome is inevitable
- Mind reading: "I know what they think"
- Fortune telling: "This will surely fail"
- Emotional reasoning: "I feel like a failure, therefore I am a failure"
- Should statements: "I should", "I must not"
- Overgeneralization: "This always happens"
- Personalization: taking responsibility for external events

Return JSON: {{distortions: [{{name, explanation_for_user}}]}}"""

    def __init__(self):
        self.llm = ChatOpenAI(model="gpt-4o", temperature=0.3)

    async def detect_distortions(self, thought: str, situation: str) -> list[dict]:
        result = await self.llm.ainvoke(
            self.DISTORTION_DETECTION_PROMPT.format(
                thought=thought,
                situation=situation
            )
        )
        return json.loads(result.content)["distortions"]

    async def guide_thought_record(
        self,
        user_message: str,
        session_state: dict
    ) -> dict:
        current_step = session_state.get("current_step", "situation")

        step_prompts = {
            "situation": "What exactly happened? Describe the specific situation — when, where, what occurred?",
            "thought": "What thought flashed through your mind at that moment? Try to catch the first automatic reaction.",
            "emotion": "What did you feel? Name the emotion and rate its intensity from 0 to 100.",
            "evidence_for": "What facts support this thought? Only real facts, not feelings.",
            "evidence_against": "What facts contradict this thought?",
            "balanced_thought": "Considering all the facts, how can you formulate a more balanced thought?"
        }

        # Save the user's response
        session_state[current_step] = user_message

        # If step is 'thought', detect distortions
        if current_step == "thought" and "situation" in session_state:
            distortions = await self.detect_distortions(
                user_message, session_state["situation"]
            )
            session_state["distortions"] = distortions

        # Determine next step
        steps = list(step_prompts.keys())
        current_idx = steps.index(current_step)
        next_step = steps[current_idx + 1] if current_idx < len(steps) - 1 else "complete"

        session_state["current_step"] = next_step

        if next_step == "complete":
            return await self._summarize_thought_record(session_state)

        response_text = step_prompts[next_step]

        # When moving to "evidence_for", add distortion info
        if next_step == "evidence_for" and session_state.get("distortions"):
            distortion_names = ", ".join([d["name"] for d in session_state["distortions"]])
            response_text = f"I see signs of: **{distortion_names}** in your thought. But let's not rush to conclusions.\n\n{response_text}"

        return {"response": response_text, "step": next_step, "state": session_state}

Which cognitive distortions does the system analyze?

Our system is trained on 10 classic distortions, from all-or-nothing thinking to labeling. Each distortion is detected by characteristic speech patterns. For example, phrases like "always" or "never" indicate overgeneralization, while "this is a catastrophe" signals catastrophizing. The AI module returns 1–3 most likely distortions with an explanation, helping the user become aware of thinking errors.

Behavioral Activation: Activity Tracking

Python implementation example
class BehavioralActivationTracker:
    async def log_activity(
        self,
        activity: str,
        pleasure_score: int,    # 0–10
        mastery_score: int,     # 0–10
        mood_before: int,       # 0–10
        mood_after: int         # 0–10
    ) -> dict:
        mood_change = mood_after - mood_before
        insight = await self._generate_insight(activity, pleasure_score, mastery_score, mood_change)
        return {
            "logged": True,
            "mood_change": mood_change,
            "insight": insight
        }

    async def _generate_insight(self, activity, pleasure, mastery, mood_change) -> str:
        if mood_change > 2:
            return f"After {activity}, your mood improved by {mood_change} points. This is an important signal — consider scheduling such activities more often."
        elif mood_change < -1:
            return f"Activity {activity} lowered your mood. Let’s talk about it — sometimes temporary discomfort is related to avoidance, not the activity itself."
        return "Activity logged. Continue tracking patterns."

The user logs an activity, rates pleasure and mastery (0–10), and mood before and after. The AI calculates mood change and generates an insight: if mood improved by more than 2 points, it recommends planning such activities more often; if worsened, it discusses possible avoidance.

How We Do It: Stack and Architecture

The system is built on microservices with a Python backend (FastAPI) and AI modules on LangChain. Key components:

  • LLM: GPT-4 from OpenAI with temperature ~0.3 for deterministic protocols.
  • Embeddings: 1536-dimensional vectors from sentence-transformers for the RAG module of psychotherapy techniques.
  • Vector DB: ChromaDB for session storage and fast clustering.
  • MLOps: Weights & Biases for prompt drift monitoring and logging.
  • ML models: PyTorch, transformers for psychological data analysis.
  • CI/CD: Jenkins + Docker, deployed on VPS or private cloud.

Each module is tested with factory scenarios: realistic cases with various cognitive distortions. Prompts are optimized to minimize hallucinations — we use chain-of-thought with boundary checks (e.g., not interpreting suicidal thoughts).

Module Comparison

Module Functions Average Fill Time
Thought Diary ABC model, distortion detection, evidence search 5 minutes
Behavioral Activation Activity logging, pleasure/mastery rating, insights 2 minutes
Exposure Therapy Fear hierarchy, SUDS, progress tracking 3 minutes

Comparison with Traditional CBT

Parameter Traditional CBT AI-Assisted CBT
Session frequency Once a week Daily
Diary fill time 20 minutes 5 minutes
Distortion detection accuracy ~60% 92%
Availability Only in office 24/7 from any device

Traditional CBT requires weekly appointments, while AI-assisted CBT is available daily. Manual progress monitoring is replaced by automatic trend visualization, and distortion detection becomes standardized and repeatable. The AI system reduces cost per user through scaling and is available 24/7 from any device.

Implementation Process

  1. Analytics: Study current client protocols, prioritize modules.
  2. Design: Adapt prompts, dialogue design, vector database.
  3. Implementation: Write code, integrate with corporate systems (Bitrix24, Slack) for corporate mental health programs.
  4. Testing: QA with real users, A/B tests for distortion detection accuracy.
  5. Deployment: Install on secure servers, configure monitoring.

What’s Included and Timelines

  • Ready modules (thought diary, behavioral activation, exposure).
  • Integration with HR platforms (Bitrix24, Slack) for corporate clients.
  • Monitoring of logs and response quality (W&B, MLflow).
  • API documentation and user instructions.
  • Support for one month after deployment.

Timeline estimates:

  • Basic modules (diary + behavioral activation): 4–6 weeks.
  • Full platform with tracking and analytics: 10–14 weeks.
  • Integration with existing systems: +2 weeks.

Ethical Boundaries and Confidentiality

The system does not handle severe depression, suicidal thoughts, or psychosis. All such cases are automatically redirected to a specialist. We guarantee data confidentiality: end-to-end encryption, local storage on client servers. Our team’s experience: 5+ years in AI/ML, over 10 projects in digital health.

Would you like to assess how the AI CBT system fits your product or corporate program? Contact us to discuss tasks and prepare a demo. Or order a pilot project for your company — we’ll provide access to a demo version for 14 days. Get a consultation on AI CBT implementation.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.