AI Personalization System for Email Newsletters in Media
Mass mailings with identical content for the entire database are an outdated approach: open rates of 15–20%, and subscriber churn reaching 5–8% per month. Our engineers observe that media outlets lose up to 70% of engagement potential by ignoring behavioral signals. We build a personalization pipeline based on LLM and event streaming that lifts open rates to 35–50%: reader profile → article scoring → subject line generation via a large language model. Technically, this is a chain involving Apache Kafka for events, Redis for profiles, and batch processing 1–2 hours before send. As a result, each subscriber receives a digest relevant to their interests, reducing churn by 2–3 times and saving up to 40% of the email marketing budget by cutting non-targeted sends.
What Data Is Critical for Personalization?
A quality profile requires at least 10–15 read articles. We collect events: views, time on page, likes, shares. Below this threshold, we use category-based personalization by selecting sections based on recent interests. Data is stored in Redis with a TTL of 30 days; topics are weighted with exponential decay. Without sufficient history, the algorithm switches to a fallback to not degrade the experience.
Why Mass Mailings Lose to Personalization?
Uniform digests ignore behavioral signals: which section a reader opens more often, which topics they skip. LLM personalization accounts for content freshness, editorial score, and even time of day. Compare:
| Metric |
Mass Mailing |
Personalized Digest |
| Open rate, % |
15–20 |
35–50 |
| Click-through rate, % |
2–5 |
8–15 |
| Monthly churn, % |
5–8 |
2–4 |
| Time spent reading |
30–60 s |
2–5 min |
After implementing personalization, open rate increased from 18% to 41% in three months, notes the technical director of one of our clients.
How We Build the Personalization Pipeline?
Step 1. Reader profiling. Collect events via Apache Kafka: views, likes, time on article. Store in Redis with TTL 30 days. Topics are weighted with exponential decay. For each reader, we form a 1536-dimensional embedding based on viewing history.
Step 2. Article scoring. For each unread article, we calculate a combined score:
- Topic match (50%)
- Freshness (30%) — newer articles score higher
- Editorial score (20%)
Step 3. Subject line generation. Use Claude 3.5 Sonnet with a few-shot prompt. The model receives the top three articles and reader interests, and outputs a headline of up to 55 characters in Russian. Example code:
from anthropic import Anthropic
def generate_personalized_digest(user_profile, available_articles, n_articles=5):
llm = Anthropic()
read_ids = user_profile.get('read_ids', set())
unread = [a for a in available_articles if a['id'] not in read_ids]
topics = user_profile.get('topics', {})
scored = []
for article in unread:
topic_score = topics.get(article.get('topic', 'general'), 0.05)
freshness = max(0, 1.0 - article.get('hours_old', 24) / 48)
quality = article.get('editorial_score', 0.7)
scored.append({**article, 'score': topic_score*0.5 + freshness*0.3 + quality*0.2})
top_articles = sorted(scored, key=lambda x: -x['score'])[:n_articles]
article_titles = [a['title'] for a in top_articles]
response = llm.messages.create(
model="claude-3-5-sonnet-20241022",
max_tokens=80,
messages=[{
"role": "user",
"content": f"Write a compelling email subject line for a news digest in Russian.\nArticles: {article_titles[:3]}\nReader's main interests: {list(topics.keys())[:3]}\nMax 55 chars. No clickbait."
}])
return {'articles': top_articles, 'subject': response.content[0].text.strip()}
The personalization trigger threshold is 10–15 read articles. Otherwise, category-based personalization (section selection) is used without article-level targeting. If needed, we fine-tune the model on a corpus of editorial materials using LoRA.
What Results Does Personalization Deliver?
Personalized digests are 2–3 times more effective than mass mailings on key metrics. We conduct A/B testing for each model, fixing lift at p99 latency under 200 ms. Our team's experience in MLOps and NLP allows rapid adaptation of the pipeline to any editorial team. Savings on non-targeted sends reach 30–40% while maintaining content quality.
| Method |
Required Data |
Latency |
Cold Start |
| Collaborative filtering |
Rating history |
High |
Problem |
| Content-based filtering |
Profiles, metadata |
Medium |
Partial |
| LLM personalization (ours) |
Behavior + NLP |
Low (batch) |
None |
What Is Included in the Turnkey Solution
- Solution architecture with block diagram
- Pipeline code in Python using LangChain and Redis
- Integration with CRM and ESP via REST API
- Documentation and team training
- 3-month warranty and support for releases
Timeline and Cost
Implementation time: from 2 to 6 weeks depending on integration complexity. Cost is calculated individually after an audit of the current stack. We will assess your project free of charge — get a consultation through the form on our website.
How to Get Started?
Tell us about your subscriber base, current metrics, and goals. We will prepare a proposal with an exact scope of work and a timeline. Contact us — we will help raise your open rate up to 50%.
Recommender System Development: From Collaborative Filtering to Real-Time Serving
On one e-commerce project with a catalog of 300k SKUs, we boosted CTR from 1.8% to 4.4% — a 2.4x increase. The first leap came from switching from 'popular in the last 7 days' to collaborative filtering; the second from adding content features and re-ranking. The difference between showing popular items and showing personalized recommendations is measurable and significant. Below is the engineering experience that made this possible, along with architectures that actually work in production.
Collaborative Filtering: Matrix Factorization and Neural Approaches
Matrix Factorization is the classic approach for implicit feedback (clicks, views, purchases without explicit ratings). ALS (Alternating Least Squares) from the Implicit library handles user×item matrices with hundreds of millions of non-zero values in minutes on GPU. Latent factors 64–256, regularization λ=0.01–0.1 are starting parameters. Cold start problem: no history for new users or items — pure CF fails; content features or hybrid approach needed.
Neural Collaborative Filtering (NCF) replaces the dot product with a neural network. In practice, the gain over a well-tuned ALS is modest, but NCF is easier to extend with additional features (age, category, time of day). Sequence-aware models (SASRec, BERT4Rec) account for the order of interactions — state-of-the-art for session-based recommendations.
How to Choose Recommender System Architecture?
The answer depends on data, load, and cold start requirements. Below are three main approaches with selection criteria.
| Criterion |
Collaborative Filtering |
Content-Based Filtering |
Hybrid (two-stage) |
| Data required |
Interaction history |
Item/user features |
Both |
| Cold start |
Poor |
Works for new items |
Partially solved |
| Diversity (long-tail) |
Low, popularity bias |
High |
Medium–High |
| Serving latency |
<5 ms (precomputed) |
<10 ms (FAISS) |
20–50 ms |
| Implementation complexity |
Low |
Medium |
High |
Hybrid architecture outperforms pure CF by 20–40% in long-tail coverage — validated on catalogs from 100k SKU.
Content-Based Filtering: When Interaction History is Scarce
Content-based recommends based on item characteristics rather than other users' behavior — solves cold start for new items. Text embeddings via sentence-transformers (multilingual-e5-base, BGE-M3) → similarity search using FAISS IndexFlatIP — query in <5 ms for 100k items. Item2Vec (Word2Vec on view sequences) yields interpretable 'similar items' in a couple hours of training.
Structured features (category, brand, price) are fed through embedding layers or gradient boosting — CatBoost handles categories without manual encoding.
Why Hybrid Models Work Better?
Production systems are almost always two-level. Stage 1 (Retrieval) — fast selection of 100–500 candidates from 300k items using ALS or Two-Tower model with vector search (FAISS, Qdrant). Stage 2 (Ranking) — heavy ranker on LightGBM or neural network with cross-features, time, device, and session context. LightFM is a good starting point for medium scale without heavy infrastructure. Our practice shows: moving from single-stage to two-stage yields a 15–25% accuracy improvement with only 20–30 ms additional latency.
Real-Time Serving: Architecture Under Load
Latency SLA — 50–100 ms at thousands of requests per second. Base recommendations precomputed (batch job hourly) → Redis by user_id → <5 ms. Real-time re-ranking via Kafka for events (clicks, cart adds) → update of context features. Feature serving — Redis with TTL (views in 24 hours, last clicked item). At 10k req/s, we deploy Redis Cluster with replication.
A/B testing is the only reliable way to measure improvements. Offline metrics do not always correlate with online. Kohavi et al., 'Online Controlled Experiments at Large Scale' (KDD 2013) — a must-read for the team. Test on 5–10% of traffic, monitor CTR, conversion, revenue per session. One of our client systems after hybridization increased revenue by 18% over a month of A/B.
Recommender System Development Timeline
The stages and typical time frames are in the table below. Costs are calculated individually based on catalog scale and latency requirements.
| Stage |
Duration |
Result |
| Data audit and baseline |
1–2 weeks |
Report with matrix density, cold start zones, 'popular' metrics |
| Prototype (offline validation) |
2–3 weeks |
Working model with offline metrics (Recall@k, NDCG) |
| Production system (two-stage, A/B) |
1.5–2.5 months |
Low-latency service with monitoring and A/B infrastructure |
| Team training and documentation |
1–2 weeks |
Model card, deployment runbook, fine-tuning session |
What's Included in Turnkey Development
- Data audit — user×item matrix density (typically <0.1%), activity distribution, temporal patterns, cold start statistics.
- Baseline — 'popular' as a simple threshold that is often hard to beat.
- Iterative improvement — ALS → content features → two-stage → sequence-aware. Each step with A/B.
- Serving infrastructure — batch precomputation, Redis, real-time re-ranking, Grafana monitoring.
- Documentation — model card with metrics, deployment instructions, feature descriptions.
- Team training — session on interpreting results and model fine-tuning.
- Support — 1 month post-launch (incident fixes, pipeline tuning).
We are a team with 7+ years of experience in recommender systems, having delivered over 30 projects for e-commerce and media. We guarantee transparent A/B testing and documented metric improvements.
Want to assess the growth potential of your catalog? Contact us for a free data audit. Order recommender system development — first prototype within two weeks.
Example ALS config for implicit feedback
from implicit.als import AlternatingLeastSquares
model = AlternatingLeastSquares(
factors=64,
regularization=0.05,
iterations=15,
use_gpu=True
)
model.fit(user_item_matrix)
More about the mathematics of recommender systems — in specialized literature.