How to Implement AI Animation of Static Photos Without Compromises?
AI photo animation is in high demand, but on-device implementation is limited: models don't fit in memory, and time-to-animation drags on. Our expertise covers high-quality portrait animation and server-side generation delivers quality but requires internet and time. We know how to combine both approaches, and with over 7 years of experience and 15+ completed mobile animation projects, our team has honed a hybrid animation approach. Architecture choice directly impacts budget: on-device saves up to 40% on GPU resources (approximately $200 per month for apps with 1000 daily users), while server-side optimizes development costs through ready-made models.
Choosing the Right Architecture for AI Animation of Static Photos
Server inference — the model lives on the backend. The app uploads the photo and receives a video. Easier to deploy, no model size constraints, can use SadTalker, LivePortrait, or AnimateDiff. Downside: needs internet, latency 3-15 seconds, GPU time cost ($0.01 to $0.05 per minute of video).
On-device — lighter specialized models. Face Reenactment via landmark-based warping (First Order Motion Model in mobile version), or simple animation via optical flow. Works offline, but quality is lower.
Most implementations choose a hybrid: on-device for quick preview (low quality), server for final result.
| Characteristic | On-device | Server |
|---|---|---|
| Quality | Medium (edge artifacts) | High (super realistic) |
| Speed | Seconds (up to 6-12 s per 1 s video) | 5-60 seconds depending on model |
| Internet | Not needed | Required |
| Usage cost | Free (after development) | GPU hours / API requests |
| Flexibility | Limited by model size | Wide model selection |
On-Device Solution Limitations
On-device animation is simple, but its quality falls short of server: noticeable artifacts, no audio sync. If you need a portrait to realistically speak, server generation is the only option. Moreover, on-device requires more time to optimize the model for a specific device: we check compatibility on 10+ iPhone and Android models.
On-Device Animation: From MediaPipe to FOMM
Lightweight approach without neural network generation: use MediaPipe Face Mesh (468 face points) to build a mesh, then deform the source image along a given motion trajectory.
// MediaPipe FaceLandmarker on iOS
let options = FaceLandmarkerOptions()
options.baseOptions.modelAssetPath = Bundle.main.path(forResource: "face_landmarker", ofType: "task")!
options.numFaces = 1
options.minFaceDetectionConfidence = 0.5
let faceLandmarker = try FaceLandmarker(options: options)
let result = try faceLandmarker.detect(image: .init(uiImage: sourcePhoto))
// landmarks.first?.faceLandmarks — 468 points [NormalizedLandmark]
// Deform via TPS (Thin Plate Spline) or affine warp
Animation — via pre-recorded head motion trajectory (mockup data) or synthetic: sinusoidal oscillations of key points with different amplitudes. Render deformed image through Metal Performance Shaders — a few milliseconds per frame.
Result — 3-5 seconds of animation, exported to .mp4 via AVAssetWriter. Quality sufficient for a "live portrait", but edge artifacts on face and background are inevitable without a full GAN.
First Order Motion Model (FOMM): Mobile Version
First Order Motion Model (FOMM) generates motion based on one driving video (donor) and a source image. On mobile runs via TFLite or ONNX Runtime, but the optimized model is 40-80 MB. On iPhone 12+, inference of one 256×256 frame: about 200-400 ms. For 30-frame animation (1 second) — 6-12 seconds processing. This is one-time generation, not real-time.
// Android: ONNX Runtime with FOMM
val session = OrtEnvironment.getEnvironment().createSession("fomm_optimized.onnx")
// Model inputs: source frame (1, 3, 256, 256) + driving frame (1, 3, 256, 256) + keypoints
val sourceInput = OnnxTensor.createTensor(env, sourceArray, longArrayOf(1, 3, 256, 256))
val drivingInput = OnnxTensor.createTensor(env, drivingArray, longArrayOf(1, 3, 256, 256))
val result = session.run(mapOf("source" to sourceInput, "driving" to drivingInput))
// Output: deformed source with applied motion
Loop over driving frames (pre-recorded motion clip): get sequence of output frames, assemble into video.
Implementing Server Generation with SadTalker and LivePortrait
For high-quality face animation with audio (talking head) — SadTalker: takes photo + audio track, generates video where the face speaks in sync with speech. On a server with A100 — 30-60 seconds per minute of video. The app uploads photo and audio, receives mp4.
LivePortrait — faster and higher quality option, 128 ms per frame on A100. API wrapper via FastAPI or Replicate. Server-based generative animation yields up to 3x more realistic results compared to on-device landmark-based methods. SadTalker is 2.5x faster than LivePortrait in per-frame inference (50 ms vs 128 ms), but LivePortrait offers higher motion realism.
// In your app, create a POST request to your server's animation endpoint with the image and optional audio as multipart data.
Polling task status or WebSocket for notification of readiness — depends on generation time.
| Model | Time per frame (A100) | Sync quality | Model size |
|---|---|---|---|
| SadTalker | ~50 ms | High | ~2 GB |
| LivePortrait | ~128 ms | Very high | ~1.5 GB |
LivePortrait outperforms SadTalker in motion realism but requires more GPU time. Choice depends on priority: speed vs quality.
How We Implement AI Animation: Process and Stages
- Requirements analysis and stack selection: define use case (on-device preview, server talking head generation, hybrid).
- Architecture design: data flow diagram, deployment model, export.
- Model implementation and integration: coding in Swift/Kotlin, server setup.
- Testing on real devices: minimum 10 iPhone and Android models.
- Deploy to App Store / Google Play with documentation.
If needed, we optimize on-device models for specific chipsets, develop our own generation pipeline, or integrate ARKit/ARCore support for overlay animation.
Export and Playback
Animation result — .mp4 (H.264 or H.265). On iOS played via AVPlayer, exported to Photos via PHPhotoLibrary. For looped animation (Living Photo) — convert to .gif via CGImageDestination or to LivePhoto format via PHLivePhoto.
Apple Live Photo: need both video file (.mov) and photo file (.jpg) with same kCGImagePropertyMakerAppleDictionary → 17 (identifier). Without this, the system Photos app does not recognize the file as a LivePhoto.
Scope of Work and Timelines
When ordering a turnkey service, you receive:
- Architectural document with model selection and justification.
- Integration of the chosen engine (MediaPipe, FOMM, SadTalker/LivePortrait).
- UI for animation style selection and trigger.
- Server part (if chosen) with task queue and statuses.
- Export to MP4/GIF/LivePhoto.
- Testing on 10+ devices with different OS versions.
- API documentation and maintenance guide.
- 3-month code warranty.
Timeline estimates: on-device landmark-based animation (single platform) — 3-4 weeks. Server integration with SadTalker/LivePortrait + both platforms — 4-7 weeks. Exact timelines depend on animation complexity and need for on-device optimization. Development costs for a mobile photo app integrating AI animation typically range from $10,000 to $50,000 depending on complexity.
Get a consultation for an accurate assessment of your project — contact us to discuss details. Order turnkey AI animation implementation, and we will select the optimal solution for your budget and timeline.







