The ad department launches 50 banners—which one will work? A/B tests require weeks and thousands of impressions. We solve this before launch: an AI system analyzes the visual features of a creative and predicts CTR with an accuracy of ±15% based on historical data. Over 5 years, we have implemented similar solutions for 30+ projects in e-commerce and fintech, guaranteeing a 10–30% CTR increase after optimization. One client, a clothing marketplace, spent a significant budget on creative testing; after implementing the AI system, they cut costs by 40% and increased average CTR by 22%. The system pays for itself in 2–3 months.
How Does AI Predict CTR?
The system extracts two types of features: semantic via CLIP and low-level visual (contrast, saturation, text coverage, presence of faces). We train LightGBM on the log of CTR to normalize the distribution. LightGBM is 3x faster than neural network alternatives with comparable accuracy.
import torch
import torch.nn.functional as F
from transformers import CLIPProcessor, CLIPModel
from PIL import Image
import numpy as np
import cv2
class CreativeFeatureExtractor:
"""
Multimodal features: CLIP semantic + low-level visual.
"""
def __init__(self):
self.clip = CLIPModel.from_pretrained(
'openai/clip-vit-large-patch14'
).eval().cuda()
self.processor = CLIPProcessor.from_pretrained(
'openai/clip-vit-large-patch14'
)
@torch.no_grad()
def extract_clip_features(
self, image: Image.Image
) -> np.ndarray:
inputs = self.processor(images=image, return_tensors='pt').to('cuda')
emb = self.clip.get_image_features(**inputs)
return F.normalize(emb, dim=-1).cpu().numpy().squeeze()
def extract_visual_features(self, image: Image.Image) -> dict:
"""
Low-level features correlating with CTR:
- face_area_ratio: face presence (face = +18% CTR per Nielsen)
- contrast: high contrast → visibility
- color_harmony: harmonious palette
- text_coverage: % of image covered by text
- brightness_variance: visual complexity
"""
img_array = np.array(image)
h, w = img_array.shape[:2]
features = {}
# Contrast (Michelson)
gray = cv2.cvtColor(img_array, cv2.COLOR_RGB2GRAY)
features['contrast_michelson'] = float(
(gray.max() - gray.min()) / (gray.max() + gray.min() + 1e-8)
)
# Brightness variance
features['brightness_variance'] = float(gray.std() / 128.0)
# Dominant colors (k=5 via k-means)
pixels = img_array.reshape(-1, 3).astype(np.float32)
criteria = (cv2.TERM_CRITERIA_EPS + cv2.TERM_CRITERIA_MAX_ITER, 20, 1.0)
_, labels, centers = cv2.kmeans(
pixels, 5, None, criteria, 5, cv2.KMEANS_PP_CENTERS
)
counts = np.bincount(labels.flatten(), minlength=5)
dominant_color = centers[counts.argmax()]
features['dominant_hue'] = float(
cv2.cvtColor(
dominant_color.reshape(1,1,3).astype(np.uint8),
cv2.COLOR_RGB2HSV
)[0,0,0] / 180.0
)
features['dominant_saturation'] = float(dominant_color.max() - dominant_color.min()) / 255.0
# Face detection
face_cascade = cv2.CascadeClassifier(
cv2.data.haarcascades + 'haarcascade_frontalface_default.xml'
)
faces = face_cascade.detectMultiScale(gray, 1.1, 4)
total_face_area = sum(fw * fh for _, _, fw, fh in faces) if len(faces) > 0 else 0
features['face_area_ratio'] = total_face_area / (h * w)
features['has_face'] = int(len(faces) > 0)
features['face_count'] = len(faces)
# Text region fraction (via morphology)
_, binary = cv2.threshold(gray, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (3, 3))
gradient = cv2.morphologyEx(binary, cv2.MORPH_GRADIENT, kernel)
features['text_coverage_estimate'] = float(
(gradient > 0).mean()
)
return features
import pandas as pd
import numpy as np
import lightgbm as lgb
from sklearn.model_selection import cross_val_score
from sklearn.preprocessing import StandardScaler
class CreativeCTRPredictor:
"""
Train on historical data: (visual_features, clip_embedding) → CTR.
Target variable: log(CTR) to normalize distribution.
"""
def __init__(self):
self.feature_extractor = CreativeFeatureExtractor()
self.model = lgb.LGBMRegressor(
n_estimators=500,
learning_rate=0.05,
max_depth=6,
min_child_samples=20,
subsample=0.8,
colsample_bytree=0.8,
reg_lambda=0.1
)
self.scaler = StandardScaler()
def build_feature_vector(self, image: Image.Image) -> np.ndarray:
visual = self.feature_extractor.extract_visual_features(image)
clip_emb = self.feature_extractor.extract_clip_features(image)
visual_vec = np.array(list(visual.values()))
# CLIP 768-dim + visual features ~10 dim
return np.concatenate([clip_emb, visual_vec])
def predict_ctr(
self, image: Image.Image
) -> dict:
features = self.build_feature_vector(image)
features_scaled = self.scaler.transform(features.reshape(1, -1))
log_ctr = self.model.predict(features_scaled)[0]
ctr_predicted = np.exp(log_ctr)
# SHAP explainability — top 3 factors
import shap
explainer = shap.TreeExplainer(self.model)
shap_values = explainer.shap_values(features_scaled)
return {
'predicted_ctr': round(float(ctr_predicted), 4),
'ctr_percentile': None, # filled from train distribution
'top_factors': self._top_shap_factors(shap_values[0], features)
}
def _top_shap_factors(
self, shap_vals: np.ndarray, feature_vals: np.ndarray, top_k: int = 3
) -> list:
top_indices = np.argsort(np.abs(shap_vals))[::-1][:top_k]
feature_names = (
[f'clip_{i}' for i in range(768)] +
['contrast', 'brightness_var', 'dominant_hue',
'dominant_sat', 'face_ratio', 'has_face',
'face_count', 'text_coverage']
)
return [
{
'feature': feature_names[i] if i < len(feature_names) else f'feat_{i}',
'shap_value': float(shap_vals[i]),
'direction': 'positive' if shap_vals[i] > 0 else 'negative'
}
for i in top_indices
]
Why LightGBM Instead of DNN?
Gradient boosting on tabular data yields better quality than neural networks with small data volumes (thousands of creatives). LightGBM is 3x faster than CatBoost on 500+ features, and SHAP interpretation works out of the box. For production, we use vLLM for CLIP inference and ONNX Runtime for model acceleration.
What Is SHAP and How Does It Help Designers?
SHAP (SHapley Additive exPlanations) assigns each feature a contribution to the prediction—similar to distributing winnings in a cooperative game. The designer sees that "face presence" increased the predicted CTR by 0.8% and "low contrast" decreased it by 0.4%. This allows conscious banner improvement rather than guessing. We integrate SHAP graphs into a Streamlit dashboard where you can upload a creative and immediately get the top three factors.
How We Integrate the System into Your Creative Workflow
We deploy a REST API on Triton Inference Server or SageMaker. The API accepts an image and returns predicted_ctr and top_factors. Through plugins for Figma and Photoshop, designers get predictions directly in the interface. Built-in MLOps infrastructure (MLflow, DVC) versions data and models, supports A/B tests of new versions. Typical integration takes 1–2 weeks.
Key Insights from Data
Based on accumulated statistics from ad platforms (Google, Meta):
| Visual Feature | Impact on CTR | Notes |
|---|---|---|
| Face presence (frontal) | +15–22% | Especially for fashion, beauty |
| High color saturation (>0.6) | +8–14% | Doesn't work for B2B |
| Text < 20% of area | +10–17% | Meta limitation 20% |
| Contrast > 0.7 | +6–11% | Visibility in feed |
| Face looking at CTA | +12% | Eye tracking studies |
| Warm colors (hue 0-60°) | +5–9% | For food, lifestyle |
Process of Work
- Analytics: audit historical creatives, collect metrics, evaluate sample size. We check that there is enough data and balance by category.
- Design: choose architecture (CLIP vs EfficientNet, LightGBM vs XGBoost), describe feature pipeline. Define quality metrics: MAE, MAPE, Spearman correlation.
- Implementation: train model, validate on holdout set (time-based split). Integrate SHAP for explainability.
- Testing: A/B experiment on 20 creatives — compare prediction with actual CTR after a week of impressions.
- Deployment: deploy API, connect to design pipeline via plugins. Set up data drift monitoring.
What's Included in the Work
- Documentation: model card, operational manual, API description.
- Source code with DVC versioning of data.
- Access to Streamlit dashboard for testing new creatives.
- Team training on the system (2 hours).
- Support for 3 months after deployment.
Timelines
| Task | Duration |
|---|---|
| Model on client's historical data (500+ creatives) | 3–5 weeks |
| Full system with API and creative workflow integration | 6–10 weeks |
| Generative optimization (AI banner correction) | 10–16 weeks |
We guarantee prediction transparency: each predicted CTR comes with the top three factors from SHAP. Get a consultation on your project—we'll evaluate your data in 2 days. Contact us to assess your data—we'll analyze 100 creatives in 2 days. Order a pilot on 20 creatives—results in a week.







