Implementing AI Video Anomaly Detection in Mobile Apps

TRUETECH is engaged in the development, support and maintenance of iOS, Android, PWA mobile applications. We have extensive experience and expertise in publishing mobile applications in popular markets like Google Play, App Store, Amazon, AppGallery and others.

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Implementing AI Video Anomaly Detection in Mobile Apps
Complex
~2-4 weeks
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    858
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    745
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1162
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1034
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    968
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    563

Implementing AI Video Anomaly Detection in Mobile Apps

We often get this request: "I want to detect falls, intrusions, or abnormal behavior on video, but I don't know where to start." The challenge is that anomalies are rare events without clear labels. It's impossible to enumerate all variations of "bad" behavior in advance. That's why we use a two-tier approach: fast deterministic rules for known scenarios and unsupervised AI for subtle anomalies. This balances speed and coverage: rules fire in milliseconds, while the AI adapts to your data.

On one project, a client wanted to monitor server room intrusions and also flag suspicious employee behavior, like unusually frequent passes outside working hours. Deterministic rules easily caught zone violations, and the autoencoder revealed atypical routes that had previously gone unnoticed. The result: 95% detection accuracy with zero false positives on normal situations.

How to Distinguish Spatial from Behavioral Anomalies?

Before writing code, we precisely define requirements with the client. Types of anomalies:

  • Spatial – object in a zone where it shouldn't be (person in server room)
  • Behavioral – normal object behaves unusually (running where others walk)
  • Temporal – event occurs at the wrong time (movement at night)
  • Technical – equipment malfunctions (smoke, vibration)

Each type requires a different architecture. For spatial anomalies, deterministic rules suffice; for behavioral, an autoencoder is needed.

How We Build the Two-Tier System

We split detection into two stages. First, fast rules based on computer vision (person detector and zones). Second, an AI model that runs only when no rule violations are present, saving resources. This approach is 3 times more CPU-efficient than using one heavy model for everything.

Tier 1: Deterministic Rules

Rules are set in configuration: forbidden zone coordinates, schedules, object types. They work without training, instantly and predictably.

class AnomalyRulesEngine {

    struct RestrictedZone {
        let polygon: [CGPoint]        // normalized coordinates
        let schedule: WorkSchedule?   // nil = always restricted
        let name: String
    }

    private let restrictedZones: [RestrictedZone]
    private let personDetector: VNCoreMLModel  // lightweight YOLOv8n

    func check(frame: CVPixelBuffer, timestamp: Date) -> [RuleViolation] {
        let persons = detectPersons(frame)
        var violations: [RuleViolation] = []

        for person in persons {
            let personCenter = person.boundingBox.center

            for zone in restrictedZones {
                if zone.polygon.contains(personCenter) {
                    if let schedule = zone.schedule, !schedule.isActive(at: timestamp) {
                        violations.append(RuleViolation(
                            type: .unauthorizedZoneAccess,
                            zone: zone.name,
                            timestamp: timestamp
                        ))
                    } else if zone.schedule == nil {
                        violations.append(RuleViolation(type: .restrictedZone, zone: zone.name))
                    }
                }
            }
        }
        return violations
    }
}

Tier 2: Autoencoder for Non-Obvious Anomalies

For behavioral anomalies, we use an autoencoder: trained on normal behavior, anomaly = high reconstruction error. The autoencoder is 10 times lighter than a YOLO detector (5 MB vs 50 MB), critical for mobile devices.

# Training autoencoder on normal video fragments
import torch
import torch.nn as nn

class VideoAnomalyAutoencoder(nn.Module):
    """
    Input tensor: [batch, frames, height, width, channels]
    Trained only on NORMAL scenes
    Anomaly: reconstruction_error > threshold
    """
    def __init__(self, input_shape=(16, 64, 64, 3)):
        super().__init__()
        self.encoder = nn.Sequential(
            nn.Conv3d(3, 32, kernel_size=(3,3,3), padding=1),
            nn.ReLU(),
            nn.MaxPool3d((1,2,2)),
            nn.Conv3d(32, 64, kernel_size=(3,3,3), padding=1),
            nn.ReLU(),
            nn.MaxPool3d((2,2,2)),
        )
        self.decoder = nn.Sequential(
            nn.ConvTranspose3d(64, 32, kernel_size=(3,3,3),
                              stride=(2,2,2), padding=1, output_padding=1),
            nn.ReLU(),
            nn.ConvTranspose3d(32, 3, kernel_size=(3,3,3),
                              stride=(1,2,2), padding=1, output_padding=(0,1,1)),
            nn.Sigmoid()
        )

    def forward(self, x):
        z = self.encoder(x)
        return self.decoder(z)

    def anomaly_score(self, x):
        reconstructed = self(x)
        return ((x - reconstructed) ** 2).mean(dim=[1,2,3,4])

The anomaly threshold is set at the 99th percentile of reconstruction error on the normal dataset. On mobile, this autoencoder is converted to CoreML or TFLite.

Mobile Inference: Sliding Window Processing

Processing each window of 16 frames takes 30-50 ms on an iPhone 12, enabling 20-30 fps.

// iOS: analyzing video stream with sliding window of 16 frames
class SlidingWindowAnalyzer {

    private var frameBuffer: CircularBuffer<CVPixelBuffer> = CircularBuffer(capacity: 16)
    private var frameCounter = 0
    private let stepSize = 8   // new window every 8 frames (50% overlap)

    func addFrame(_ frame: CVPixelBuffer) async -> AnomalyScore? {
        frameBuffer.append(frame)
        frameCounter += 1

        guard frameCounter % stepSize == 0,
              frameBuffer.count == 16 else { return nil }

        return try? await computeAnomalyScore(frames: Array(frameBuffer))
    }

    private func computeAnomalyScore(frames: [CVPixelBuffer]) async throws -> AnomalyScore {
        let tensor = prepareTensor(frames)  // [1, 16, 64, 64, 3]
        let output = try autoencoderModel.prediction(input: tensor)
        let score = output.anomalyScore.floatValue

        return AnomalyScore(
            value: score,
            isAnomaly: score > anomalyThreshold,
            frameWindow: frames
        )
    }
}

Alerts and Reaction

We implement a multi-level system: warnings and critical alerts with cooldown to avoid spam. Critical alerts go via webhook to an external security system.

// Android: multi-level alert system
sealed class AnomalyAlert {
    data class Warning(val message: String, val score: Float) : AnomalyAlert()
    data class Critical(val message: String, val violations: List<RuleViolation>) : AnomalyAlert()
}

class AlertManager(private val notificationManager: NotificationManager) {

    private val cooldownMap = mutableMapOf<String, Long>()
    private val alertCooldownMs = 30_000L  // no spam: max once per 30 sec

    fun emit(alert: AnomalyAlert, alertKey: String) {
        val lastAlertTime = cooldownMap[alertKey] ?: 0L
        if (System.currentTimeMillis() - lastAlertTime < alertCooldownMs) return

        cooldownMap[alertKey] = System.currentTimeMillis()

        when (alert) {
            is AnomalyAlert.Warning -> showLocalNotification(alert.message, priority = LOW)
            is AnomalyAlert.Critical -> {
                showLocalNotification(alert.message, priority = HIGH)
                sendWebhook(alert)  // integration with external security system
            }
        }
    }
}

Comparison of Approaches

Feature Deterministic Rules Autoencoder AI
Anomaly type Spatial, temporal Behavioral, technical
Training Expert-defined zones/schedule Unsupervised on normal data
Accuracy 100% for defined rules Data-dependent (ROC-AUC ~0.9)
Model size 0 (lightweight face detector) 5-15 MB
Latency on iOS <5 ms 30-50 ms per window

Deterministic rules are 10 times faster than AI, but miss new anomaly types. The autoencoder finds what wasn't predefined — perfect as a second tier.

What Our Work Includes

Our team (8+ years in mobile CV, 15+ projects) offers:

  • Collect and label normal data for training
  • Design deterministic rules for your scenario
  • Train and convert autoencoder to CoreML / TFLite
  • Integrate with iOS (Swift 5.9+, SwiftUI) and Android (Kotlin, Jetpack Compose)
  • Multi-level alert system with webhook integration
  • Documentation and operator training
  • Post-deployment support

We'll assess your project within 2 days. Contact us for a consultation — we'll show a demo on real data. Request a pilot launch and see the effectiveness.

Timeline Estimates

Stage Duration
Zone violation detection (without AI) 1-2 weeks
Full system with autoencoder, sliding window, alerts 2-4 weeks
Collect normal data and train model +1-2 weeks
Integration with external security system +1 week

Cost is calculated individually. We guarantee code quality and full documentation. We use CoreML and TFLite — proven tools for mobile inference.

Machine Learning in Mobile Apps: CoreML, TFLite, and On-Device Models

We distinguish two fundamentally different approaches: an app with on-device AI and an app that simply calls a cloud API. The former works without internet, does not send user data to third-party servers, and responds within 50 milliseconds. The latter depends on network latency and pricing plans. Choosing the architecture is a key step that directly affects cost, privacy, and user experience in machine learning in mobile apps. Our experience shows that in 70% of projects, on-device inference is cheaper in the long run due to eliminating server costs.

How to Choose Between CoreML and TFLite for On-Device Inference?

CoreML — Apple's native framework for running ML models on device. Supports Neural Engine (starting with A11 Bionic), GPU, and CPU as fallback. Models are converted to .mlmodel format via coremltools from PyTorch, ONNX, or TensorFlow. Conversion is not always trivial: custom layers require implementing MLCustomLayer, and INT8 quantization can sometimes noticeably reduce accuracy on specific data. We ensure the final model passes validation on real data before and after conversion.

TensorFlow Lite — cross-platform alternative for Android and Flutter. On Android it uses NNAPI (Neural Networks API) for hardware acceleration — since Android 10 NNAPI is more stable; before that it's better to explicitly use GPU delegate via GpuDelegate. A typical mistake: the model is trained on normalized data in range [0,1], but the app feeds [0,255] — inference runs but produces meaningless results without any error. We include an automatic input data validation module in the SDK.

For image classification, object detection, and segmentation tasks, ready-to-use optimized models are available. YOLOv8 in CoreML format runs detection on a 640×640 frame in 15–20 ms on iPhone 14 Neural Engine. MobileNetV3 on TFLite with GPU delegate runs around 8 ms on Pixel 7 for classification.

Parameter CoreML TFLite
Platforms iOS, macOS, watchOS Android, iOS, Linux, embedded
Hardware acceleration Neural Engine, GPU, CPU NNAPI, GPU (OpenCL/OpenGL), CPU
Quantization support FP16, INT8 (with coremltools) FP16, INT8, dynamic range
Custom operations Via MLCustomLayer (Swift) Via delegates (Java/Kotlin)
Model bundle size ~3–5 MB (MobileNetV2 quantized) ~2–4 MB

What If You Need Text Generation On-Device?

Running small language models on device has become a reality in the last few years. Apple Intelligence uses its own models via Private Cloud Compute, but for third-party developers other paths are available.

llama.cpp with Metal backend on iOS is a working approach for phi-3-mini (3.8B parameters, 4-bit quantization, ~2.3 GB). Inference: 15–25 tokens/second on iPhone 15 Pro. For integration in Swift, use the Swift Package llama.swift or a wrapper via C interface llama.h. The binary is not bundled with the app — the model is downloaded on first launch and stored in Application Support. Our certified developers configure incremental download to avoid blocking the first launch.

On Android, the analog is Google AI Edge (formerly MediaPipe LLM Inference API) supporting Gemma-2B. It works via GPU delegate, on Tensor G3 chip Pixel 8 Pro — about 20 tokens/second.

Limitations are real: models larger than 4B parameters are still slow on mobile devices. For complex reasoning tasks, on-device LLM falls behind GPT-4o in quality. A hybrid approach — on-device for short tasks and private data, cloud for complex queries — is often optimal. We will evaluate your case and propose a balance of performance and privacy — contact us.

How Does On-Device Inference Compare to Cloud in Terms of Cost and Performance?

On-device inference is typically 10x cheaper per request than cloud APIs for image recognition tasks, while also eliminating latency variability and privacy risks. The table below summarizes the trade-offs.

Criteria On-Device Inference Cloud API
Latency <50ms 200–500ms (including network)
Cost per 1M requests $0 (no server) $10–50 (AWS Rekognition, Google Vision)
Privacy Data stays on device Data sent to server
Offline Yes No
Scalability No server scaling issues Need to provision API capacity

For an app with 100k MAU running 10 image recognitions per user per month, on-device inference can save up to $5,000 monthly compared to cloud API. Get a free consultation on your ML architecture today.

Integrating OpenAI API and Other Cloud Models

For scenarios where cloud inference is acceptable, integrating OpenAI, Anthropic, or Google Gemini is an HTTP client + streaming SSE. In Swift, AsyncThrowingStream is convenient for streaming responses. In Kotlin, use Flow.

Critically: API keys must never be stored in the app bundle. Even an obfuscated key can be extracted from the IPA in 10 minutes using strings or frida. Correct architecture: mobile app → your own backend → OpenAI API. The backend controls rate limiting, logs requests, and protects the key.

What Is Included in the Work (Deliverables)

  • Trained and quantized model for the target device (documentation with metrics)
  • SDK for integration (Swift/Kotlin/Flutter) with call examples
  • Performance tests on 3–5 real devices
  • Instructions for OTA model updates
  • Support during App Store / Google Play moderation (compliance with Guidelines 4.2, 5.1)
  • 2 weeks of technical support after release

Typical Project Pipeline

  1. Task analysis — measure latency, privacy, size, supported devices.
  2. Model prototyping — in Python, evaluate accuracy on target data.
  3. Conversion and quantization — for CoreML/TFLite with validation.
  4. Integration into the app — model wrapped in a service layer (easy to swap CoreML ↔ TFLite ↔ cloud).
  5. Testing — on real devices, measure FPS, RAM, battery.
  6. Deployment — via TestFlight / Firebase App Distribution, monitor metrics.

Timelines: integration of a ready CoreML/TFLite model — 1–2 weeks, development of a custom model with mobile optimization — from 6 weeks, on-device LLM chat with personalization — 4–8 weeks.

Why We Take on Complex Cases?

10+ years of experience in mobile development, 50+ implemented AI/ML solutions, guarantee of compatibility with current iOS and Android versions. All projects undergo code review and load testing. The cost includes preparation of moderation documentation and training of your team.

Contact us — we will help you choose the architecture and implement ML in your app turnkey. Order an audit of your existing solution — we will assess the potential for server cost savings free of charge. In some projects, savings can reach significant amounts per month.