Behavior Detection System Development
A duty operator staring at 16 monitors with 4 cameras each—that's classic. After 20 minutes, attention fades; after an hour, they will miss a real fight. Automatic detection of unwanted behavior solves this problem. We develop systems that see not only 'what' but also 'how': poses, trajectories, context. The difficulty is that the boundary between normal and abnormal is blurred, and high recall requires addressing false positives. The system is built modularly: rule-based detectors for simple scenarios, skeleton analysis for falls and running, deep neural networks for aggression and vandalism. Each module cascades filters to keep precision at 82% and above. Assess automation possibilities for your facility—contact us for a preliminary analysis.
What Behavior Types Does the System Recognize?
Level 1 (rule-based): simple events—line crossing, point accumulation. No ML needed, low CPU. Configured in a day.
Level 2 (pose-based): events based on a human skeleton (MediaPipe / RTMPose). Analysis of joint angles and movement speed. Delivers 90% accuracy on falls without GPU.
Level 3 (video-based): deep video understanding via 3D CNN or Video Transformer. High accuracy, requires GPU. We use the KINETICS-400 dataset for pretraining.
How Does Aggression Detection Work?
A fight is not just two objects in frame—it's characteristic dynamics: sharp arm movements, sudden falls, increased movement speed. We use a 3D CNN trained on 400 classes from KINETICS. The model 'watches' 16 frames (0.5 s) and outputs aggression probability. In practice, precision=82%, recall=88%—better than 90% of operators.
import torch
import torch.nn as nn
from torchvision.models.video import r3d_18, R3D_18_Weights
class FightDetector:
def __init__(self, model_path: str, threshold: float = 0.7):
base = r3d_18(weights=R3D_18_Weights.KINETICS400_V1)
base.fc = nn.Sequential(
nn.Linear(512, 128),
nn.GELU(),
nn.Dropout(0.4),
nn.Linear(128, 2) # fight / no_fight
)
base.load_state_dict(torch.load(model_path))
base.eval()
self.model = base
self.threshold = threshold
# Sliding video window
self.frame_buffer = []
self.window_size = 16 # 16 frames = ~0.5 sec at 30fps
def update(self, frame: np.ndarray) -> dict | None:
"""Update buffer and get result"""
self.frame_buffer.append(frame)
if len(self.frame_buffer) > self.window_size:
self.frame_buffer.pop(0)
if len(self.frame_buffer) == self.window_size:
return self._classify_window()
return None
@torch.no_grad()
def _classify_window(self) -> dict:
# [T, H, W, C] → [1, C, T, H, W]
clip = np.stack(self.frame_buffer)
clip = torch.from_numpy(clip).float().permute(3, 0, 1, 2)
clip = self._normalize(clip).unsqueeze(0)
logits = self.model(clip)
probs = torch.softmax(logits, dim=1).squeeze()
fight_prob = float(probs[1])
return {
'fight_detected': fight_prob > self.threshold,
'confidence': fight_prob
}
What About False Positives?
The main pain of aggression detectors is false positives. Solution: temporal confirmation (N consecutive frames)—cuts 70% of random false positives; multi-evidence fusion (skeleton + video + context)—reduces errors in rooms with glare; human-in-the-loop—sends uncertain cases to the guard, and confirmed ones go to retraining. Ultimately we achieve precision=82% with recall=88% on fights. This saves up to 30% of the guard budget by reducing false alarms.
Technical detail: cascade filtering
First level—rule-based (motion trigger), second—pose analysis (posture check), third—3D CNN (event verification). Each level filters out false positives, leaving only confirmed incidents.Skeleton-based Behavior Analysis
Skeleton analysis enables detection of falls, running, suspicious lingering without costly 3D CNN. It only needs 5–15 FPS and a lightweight neural network. The code below processes trackers over 30 frames and returns the dominant behavior.
import numpy as np
from collections import deque
class BehaviorAnalyzer:
def __init__(self, window_size: int = 30):
self.track_history = {} # track_id -> deque of (frame, keypoints)
self.window = window_size
def update(self, track_id: int, frame_num: int,
keypoints: dict) -> dict:
if track_id not in self.track_history:
self.track_history[track_id] = deque(maxlen=self.window)
self.track_history[track_id].append((frame_num, keypoints))
if len(self.track_history[track_id]) < 10:
return {'behavior': 'unknown'}
return self._analyze(track_id)
def _analyze(self, track_id: int) -> dict:
history = list(self.track_history[track_id])
keypoints_seq = [kp for _, kp in history]
behaviors = {
'fall': self._detect_fall(keypoints_seq),
'running': self._detect_running(keypoints_seq),
'crouching': self._detect_crouching(keypoints_seq[-1]),
'loitering': self._detect_loitering(keypoints_seq)
}
dominant = max(behaviors, key=lambda k: behaviors[k])
return {
'behavior': dominant if behaviors[dominant] > 0.5 else 'normal',
'scores': behaviors
}
def _detect_fall(self, seq: list) -> float:
"""Fall detection: sharp drop of center of mass"""
hip_ys = []
for kp in seq:
if kp.get('left_hip') and kp.get('right_hip'):
avg_hip_y = (kp['left_hip']['y'] + kp['right_hip']['y']) / 2
hip_ys.append(avg_hip_y)
if len(hip_ys) < 10:
return 0.0
# Normalized coordinates: y increases downwards
max_drop = max(hip_ys[-5:]) - min(hip_ys[-15:-5]) if len(hip_ys) >= 15 else 0
return min(1.0, max_drop / 0.3) # 0.3 = 30% of frame height
def _detect_loitering(self, seq: list) -> float:
"""Loitering detection: person stays in one place for a long time"""
if len(seq) < 20:
return 0.0
positions = [(kp.get('nose', {}).get('x', 0.5),
kp.get('nose', {}).get('y', 0.5))
for kp in seq]
positions = np.array(positions)
spread = np.std(positions, axis=0).mean()
return min(1.0, (0.05 - spread) / 0.05) # < 5% spread = loitering
Rule-based vs Deep Learning: When to Choose What
For counters and line crossing, rule-based is faster and cheaper. For fights and vandalism, only deep learning works. Skeleton analysis (pose) is the sweet spot: gives 90% accuracy on falls without heavy GPU. In our practice, combining rule-based + pose saves up to 50% of the computing budget.
What's Included in the Work
- Documentation: scenario specification, architecture diagram, alert API description.
- Deliverables: Docker images with models, config repository, deployment instructions.
- Integration: adaptation to VMS (Milestone, Genetec, TRASSIR) or RTSP streams.
- Training: a session for operators and administrators (up to 4 hours).
- Support: 3 months of warranty maintenance, model updates upon retraining.
Turnkey Implementation Process
- Analysis of scenarios and zones at the site (1–2 days).
- Collection and labeling of data (if retraining needed).
- Architecture selection: rule-based, pose, 3D CNN, or hybrid.
- Integration with existing video surveillance system.
- Cascade filtering and sensitivity tuning.
- Testing on historical recordings and launch 24/7.
Timelines start from 4 weeks for a basic solution. Order a pilot on two cameras and see results in two weeks.
| Behavior Type | Precision | Recall |
|---|---|---|
| Fall | 91% | 94% |
| Running/Rushing | 88% | 92% |
| Aggression/Fight | 82% | 88% |
| Vandalism | 79% | 83% |
| Pickpocketing | 74% | 79% |
| Scale | Timeline |
|---|---|
| 2–3 event types, rule-based + pose | 4–6 weeks |
| Full behavior analytics | 9–14 weeks |
| High-precision system with training | 14–22 weeks |
Over 50 deployments in retail, logistics, and offices—our team has 5+ years of experience in Computer Vision. We'll assess your project free of charge; reach out to us.







