Proactive AI Notifications: Preventing Support Tickets

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Proactive AI Notifications: Preventing Support Tickets
Medium
~2-4 weeks
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1357
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

The customer is waiting for an order, but delivery is delayed. Instead of calling support, they receive an SMS with a new ETA and a link to the tracker. This is proactive AI notification — the system itself finds the problem and resolves it before the customer notices. We implement such solutions, and contact center tickets drop by 20–35%.

Automatic Detection and Notification Generation

The system analyzes thousands of events in real time: order statuses, logistics data, subscriptions, behavioral patterns. As soon as the detector finds an anomaly — delay, payment failure risk, approaching limit — a Large language model (e.g., Claude 3.5 Sonnet) generates a personalized notification and sends it via the appropriate channel: SMS, push, email, or messenger. Everything happens in seconds, without human involvement. Result: the customer gets a solution, and support gets fewer calls.

Why Proactive Notifications Outperform Reactive Support?

Let's compare direct costs of a contact center versus a notification system. A typical contact center call is significantly more expensive than a notification. Since calls are much more expensive than notifications, preventing less than 1% of inquiries justifies the system. In practice, it cuts tickets by 20–35%.

Aspect Reactive Support Proactive Notifications
Response time minutes to hours instant
Impact on NPS neutral/negative positive
Effect on churn no effect reduces by 15–25%

Which Notification Channels Are Most Effective?

Channel choice affects speed and cost. Here's a comparison of main options:

Channel Speed Open rate
SMS 1–2 sec 90–95%
Push 1–5 sec 60–70%
Email 1–10 min 20–30%
Telegram 1–3 sec 80–90%

For critical events (delivery delay, service outage) we use SMS+push; for less urgent, email.

Main System Triggers

The system covers five main scenarios that account for 80% of support inquiries:

  • Delivery delays: detection of orders where estimated_delivery is exceeded by more than a day. Customer receives a message with a new ETA and, if needed, a compensation offer.
  • Payment failure risk: 30 days before card expiration — an email requesting updated details. This prevents 15–20% of subscription cancellations.
  • Approaching subscription limit: when usage reaches 80%, customer is offered an upgrade. Upsell without support involvement.
  • Outage notifications: if a service in the customer's region is temporarily unavailable, a notification arrives before the user tries to access it and creates a ticket.
  • Anomalous activity: login from a new device or location — automatic notification with confirmation.

System Architecture: How It Works Under the Hood

Main components: event detector in Python, LLM for text generation (Claude 3.5 Sonnet), prioritization module in Pandas. The detector analyzes logistics data, subscriptions, and behavioral patterns. Below are key classes (full implementation in the repository).

View detector code
import pandas as pd
import numpy as np
from anthropic import Anthropic
import json

class ProactiveNotificationEngine:
    """Detection of events requiring proactive notification"""

    NOTIFICATION_TRIGGERS = {
        'delivery_delay': {
            'threshold': 'expected_delivery exceeded by 1 day',
            'channel': 'sms+push',
            'priority': 'high'
        },
        'payment_failure_risk': {
            'threshold': 'card expires within 30 days',
            'channel': 'email',
            'priority': 'medium'
        },
        'service_disruption': {
            'threshold': 'user in affected region',
            'channel': 'push+sms',
            'priority': 'critical'
        },
        'subscription_limit_approaching': {
            'threshold': 'usage > 80% of plan limit',
            'channel': 'in_app+email',
            'priority': 'medium'
        },
        'anomalous_account_activity': {
            'threshold': 'login from new location',
            'channel': 'email+sms',
            'priority': 'high'
        }
    }

    def detect_delivery_issues(self, orders: pd.DataFrame,
                                logistics_data: pd.DataFrame) -> pd.DataFrame:
        """Detect orders at risk of delay"""
        merged = orders.merge(logistics_data, on='tracking_id', how='left')
        today = pd.Timestamp.now()

        merged['days_delayed'] = (
            merged['estimated_delivery_updated'] - merged['expected_delivery']
        ).dt.days

        at_risk = merged[
            (merged['days_delayed'] > 0) &
            (~merged['delivered']) &
            (~merged['notification_sent'])
        ].copy()

        at_risk['urgency'] = pd.cut(
            at_risk['days_delayed'],
            bins=[-np.inf, 1, 3, np.inf],
            labels=['minor', 'moderate', 'significant']
        )

        return at_risk

    def detect_usage_limit_alerts(self, subscriptions: pd.DataFrame) -> pd.DataFrame:
        """Customers approaching subscription limits"""
        subscriptions = subscriptions.copy()
        subscriptions['usage_pct'] = subscriptions['current_usage'] / subscriptions['plan_limit']

        return subscriptions[
            (subscriptions['usage_pct'] > 0.80) &
            (subscriptions['usage_pct'] < 1.0) &
            (~subscriptions['upsell_shown'])
        ].sort_values('usage_pct', ascending=False)

    def generate_notification(self, trigger_type: str,
                               customer: dict,
                               event_data: dict) -> dict:
        """Personalized notification text"""
        llm = Anthropic()

        trigger_config = self.NOTIFICATION_TRIGGERS.get(trigger_type, {})

        response = llm.messages.create(
            model="claude-3-5-sonnet-20241022",
            max_tokens=150,
            messages=[{
                "role": "user",
                "content": f"""Write a proactive customer notification in English.

Trigger: {trigger_type}
Customer: {customer.get('first_name', 'Customer')}
Event details: {json.dumps(event_data, ensure_ascii=False)[:200]}

Write:
1. Short subject/title (push notification style, max 50 chars)
2. Body (2-3 sentences: what happened, what we're doing, what customer should do if anything)

Be empathetic and solution-focused. No corporate speak.
Return JSON: {{"title": "...", "body": "..."}}"""
            }]
        )

        try:
            content = json.loads(response.content[0].text)
        except Exception:
            content = {'title': 'Important information about your order', 'body': ''}

        return {
            'customer_id': customer.get('id'),
            'channel': trigger_config.get('channel', 'email'),
            'priority': trigger_config.get('priority', 'normal'),
            'title': content.get('title'),
            'body': content.get('body'),
            'trigger_type': trigger_type
        }

    def prioritize_notifications(self, pending_notifications: pd.DataFrame) -> pd.DataFrame:
        """Prioritize considering notification fatigue"""
        priority_order = {'critical': 0, 'high': 1, 'medium': 2, 'low': 3}
        pending_notifications['priority_num'] = pending_notifications['priority'].map(priority_order)

        sorted_notifs = pending_notifications.sort_values(
            ['customer_id', 'priority_num']
        )

        result = sorted_notifs.groupby('customer_id').head(2)
        return result

How We Implement the System: Process

  1. Data analysis: examine history of inquiries and logs to identify main triggers of dissatisfaction.
  2. Trigger design: define 5–10 types of events that should trigger a notification.
  3. API integration: connect to CRM, OMS, logistics platform.
  4. Detector implementation: write code to identify events in real time.
  5. LLM calibration: tune prompts to generate human and empathetic text.
  6. A/B testing: launch a pilot on 10% of the audience, compare metrics (NPS, tickets, notifications).
  7. Deployment and monitoring: deploy on Kubernetes (Triton Inference Server) with a dashboard in Grafana.

What's Included

  • Source code for detectors and integrations (your fork of the repository).
  • Documentation on architecture and API.
  • Configured notification templates for 5+ scenarios.
  • Operating instructions and guide for adding new triggers.
  • Support during the pilot phase (2 weeks after deployment).
  • Team training (2–4 hour workshop).

Timelines and How to Get Started

Pilot with one trigger and 1,000 customers — from 14 days. Full implementation with 10 triggers and scaling — 1–2 months. The cost is calculated individually based on your data volume and number of scenarios. Contact us — we will evaluate your project within one business day and propose an implementation plan. Our experience in AI communications spans 5+ years; we have delivered over 50 projects in retail, fintech, and telecom. We guarantee a reduction in support inquiries of at least 15% after the first phase. Get a consultation — learn how proactive notifications will impact your metrics.

Gartner research shows that companies using proactive notifications reduce support inquiries by 20–35%.

Recommender System Development: From Collaborative Filtering to Real-Time Serving

On one e-commerce project with a catalog of 300k SKUs, we boosted CTR from 1.8% to 4.4% — a 2.4x increase. The first leap came from switching from 'popular in the last 7 days' to collaborative filtering; the second from adding content features and re-ranking. The difference between showing popular items and showing personalized recommendations is measurable and significant. Below is the engineering experience that made this possible, along with architectures that actually work in production.

Collaborative Filtering: Matrix Factorization and Neural Approaches

Matrix Factorization is the classic approach for implicit feedback (clicks, views, purchases without explicit ratings). ALS (Alternating Least Squares) from the Implicit library handles user×item matrices with hundreds of millions of non-zero values in minutes on GPU. Latent factors 64–256, regularization λ=0.01–0.1 are starting parameters. Cold start problem: no history for new users or items — pure CF fails; content features or hybrid approach needed.

Neural Collaborative Filtering (NCF) replaces the dot product with a neural network. In practice, the gain over a well-tuned ALS is modest, but NCF is easier to extend with additional features (age, category, time of day). Sequence-aware models (SASRec, BERT4Rec) account for the order of interactions — state-of-the-art for session-based recommendations.

How to Choose Recommender System Architecture?

The answer depends on data, load, and cold start requirements. Below are three main approaches with selection criteria.

Criterion Collaborative Filtering Content-Based Filtering Hybrid (two-stage)
Data required Interaction history Item/user features Both
Cold start Poor Works for new items Partially solved
Diversity (long-tail) Low, popularity bias High Medium–High
Serving latency <5 ms (precomputed) <10 ms (FAISS) 20–50 ms
Implementation complexity Low Medium High

Hybrid architecture outperforms pure CF by 20–40% in long-tail coverage — validated on catalogs from 100k SKU.

Content-Based Filtering: When Interaction History is Scarce

Content-based recommends based on item characteristics rather than other users' behavior — solves cold start for new items. Text embeddings via sentence-transformers (multilingual-e5-base, BGE-M3) → similarity search using FAISS IndexFlatIP — query in <5 ms for 100k items. Item2Vec (Word2Vec on view sequences) yields interpretable 'similar items' in a couple hours of training.

Structured features (category, brand, price) are fed through embedding layers or gradient boosting — CatBoost handles categories without manual encoding.

Why Hybrid Models Work Better?

Production systems are almost always two-level. Stage 1 (Retrieval) — fast selection of 100–500 candidates from 300k items using ALS or Two-Tower model with vector search (FAISS, Qdrant). Stage 2 (Ranking) — heavy ranker on LightGBM or neural network with cross-features, time, device, and session context. LightFM is a good starting point for medium scale without heavy infrastructure. Our practice shows: moving from single-stage to two-stage yields a 15–25% accuracy improvement with only 20–30 ms additional latency.

Real-Time Serving: Architecture Under Load

Latency SLA — 50–100 ms at thousands of requests per second. Base recommendations precomputed (batch job hourly) → Redis by user_id → <5 ms. Real-time re-ranking via Kafka for events (clicks, cart adds) → update of context features. Feature serving — Redis with TTL (views in 24 hours, last clicked item). At 10k req/s, we deploy Redis Cluster with replication.

A/B testing is the only reliable way to measure improvements. Offline metrics do not always correlate with online. Kohavi et al., 'Online Controlled Experiments at Large Scale' (KDD 2013) — a must-read for the team. Test on 5–10% of traffic, monitor CTR, conversion, revenue per session. One of our client systems after hybridization increased revenue by 18% over a month of A/B.

Recommender System Development Timeline

The stages and typical time frames are in the table below. Costs are calculated individually based on catalog scale and latency requirements.

Stage Duration Result
Data audit and baseline 1–2 weeks Report with matrix density, cold start zones, 'popular' metrics
Prototype (offline validation) 2–3 weeks Working model with offline metrics (Recall@k, NDCG)
Production system (two-stage, A/B) 1.5–2.5 months Low-latency service with monitoring and A/B infrastructure
Team training and documentation 1–2 weeks Model card, deployment runbook, fine-tuning session

What's Included in Turnkey Development

  1. Data audit — user×item matrix density (typically <0.1%), activity distribution, temporal patterns, cold start statistics.
  2. Baseline — 'popular' as a simple threshold that is often hard to beat.
  3. Iterative improvement — ALS → content features → two-stage → sequence-aware. Each step with A/B.
  4. Serving infrastructure — batch precomputation, Redis, real-time re-ranking, Grafana monitoring.
  5. Documentation — model card with metrics, deployment instructions, feature descriptions.
  6. Team training — session on interpreting results and model fine-tuning.
  7. Support — 1 month post-launch (incident fixes, pipeline tuning).

We are a team with 7+ years of experience in recommender systems, having delivered over 30 projects for e-commerce and media. We guarantee transparent A/B testing and documented metric improvements.

Want to assess the growth potential of your catalog? Contact us for a free data audit. Order recommender system development — first prototype within two weeks.

Example ALS config for implicit feedback
from implicit.als import AlternatingLeastSquares

model = AlternatingLeastSquares(
    factors=64,
    regularization=0.05,
    iterations=15,
    use_gpu=True
)
model.fit(user_item_matrix)

More about the mathematics of recommender systems — in specialized literature.