Creating a 3D model from photos directly on a mobile device — a task that seemed like science fiction not long ago. Today, thanks to NeRF and 3D Gaussian Splatting, it's a reality. Most photogrammetry solutions require powerful desktop GPUs and are not adapted for mobile UX — we close this gap by offering a ready-made pipeline with guided capture and cloud processing. We develop such turnkey solutions for iOS and Android. Our experience: 5+ years in mobile development and 30+ projects in computer vision. We guarantee reconstruction accuracy and full support at all stages. Order development now — get a solution ready for publication.
What problem do we solve?
Manual 3D modeling takes hours and requires skills. Automatic reconstruction often produces artifacts due to poor coverage or low image quality. We eliminate these problems with guided capture and an optimized pipeline based on NeRF and 3D Gaussian Splatting. The user simply walks around the object, following prompts, and our system does the rest. Scenes with reflective or uniform surfaces are particularly challenging — for these we use an adaptive capture strategy.
What to choose: NeRF, Gaussian Splatting, or photogrammetry?
The three technologies solve the same task: 3D object from photos. The difference is fundamental:
| Method | Training Speed | Render Time | Quality | Editability |
|---|---|---|---|---|
| Classic NeRF | Hours–days | Slow | High | Poor |
| InstantNGP/Nerfacto | 5–30 min | Fast | Good | Fair |
| 3D Gaussian Splatting | 10–40 min | Real-time | Excellent | Good |
| Photogrammetry (Metashape, COLMAP) | 30 min–several hours | Instant (mesh) | Depends on photos | Excellent |
For mobile applications, 3D Gaussian Splatting is currently the best balance of speed and quality. It is 3–5 times faster than classic NeRF with comparable quality. For quick AR previews, photogrammetry with a modern COLMAP pipeline is suitable.
How does 3D reconstruction on mobile work?
On-device reconstruction is only possible in limited scenarios (e.g., Apple Object Capture API — only on Mac with Apple Silicon). A practical architecture for mobile:
- Guided capture — guided capture with AR overlay (ARKit/ARCore) collects 20–60 photos and camera pose metadata.
- On-device validation — check coverage, sharpness, and frame count.
- Upload to cloud — compressed images and metadata are sent to a GPU instance.
- Point cloud building — COLMAP SfM (if no poses) or metadata import.
- Training 3D Gaussian Splatting — 10–40 minutes on T4/A100.
- Export — .glb for AR Quick Look, .splat for web viewer.
- AR viewing — load model and render via RealityKit/SceneViewer.
Technical requirements for cloud GPU
For training Gaussian Splatting, an NVIDIA GPU with 8+ GB VRAM (T4, A10G, A100) is recommended. Training time ranges from 10 to 40 minutes depending on the number of frames and resolution. We use containerization (Docker + NVIDIA Container Toolkit).Guided capture on iOS with ARKit
Key UX: the user must walk around the object correctly, otherwise reconstruction will have artifacts.
class GuidedCaptureSession: NSObject {
private var arSession: ARSession
private var capturedFrames: [(UIImage, simd_float4x4)] = [] // image + camera transform
private let targetFrameCount = 40
private let minAngleBetweenFrames: Float = 8.0 // degrees
func shouldCaptureFrame(currentTransform: simd_float4x4) -> Bool {
guard let lastTransform = capturedFrames.last?.1 else { return true }
// Angular distance from the last captured frame
let angularDistance = computeAngularDistance(currentTransform, lastTransform)
return angularDistance >= minAngleBetweenFrames
}
var captureProgress: Float {
// Estimate orbit coverage around the object
let coveredAngles = estimateOrbitCoverage(capturedFrames.map { $0.1 })
return min(coveredAngles / 360.0, 1.0)
}
}
The AR overlay shows an "orbit" around the object: green arcs — already captured angles, gray — need to be captured. This reduces the rate of failed reconstructions due to incomplete coverage.
Image quality requirements for AI 3D reconstruction
Before sending to the cloud, basic validation is performed on the device:
func validateCaptureSet(_ frames: [(UIImage, simd_float4x4)]) -> ValidationResult {
// Minimum number of frames
guard frames.count >= 20 else {
return .insufficientFrames(current: frames.count, required: 20)
}
// Angle coverage (need at least 270° out of 360°)
let orbitCoverage = estimateOrbitCoverage(frames.map { $0.1 })
guard orbitCoverage >= 0.75 else {
return .insufficientCoverage(coverage: orbitCoverage)
}
// Average frame sharpness
let avgSharpness = frames.map { sharpnessScore($0.0) }.reduce(0, +) / Float(frames.count)
guard avgSharpness >= 60.0 else {
return .blurryImages
}
return .valid
}
Backend: Training 3D Gaussian Splatting
ARKit metadata (camera poses) simplifies COLMAP SfM, reducing processing time. If poses are missing, we run SfM from nerfstudio.
# nerfstudio + gsplat pipeline
from nerfstudio.cameras.cameras import CameraType
from nerfstudio.pipelines.base_pipeline import Pipeline
def run_gaussian_splatting(
images_dir: Path,
camera_poses: list[np.ndarray] | None = None,
output_dir: Path = Path("output")
) -> Path:
"""
If camera_poses are provided (from ARKit) — skip COLMAP SfM.
This reduces processing time from 15-20 minutes to 3-5 minutes.
"""
config = SplatfactoModelConfig(
num_downscales=2, # reduce for speed
use_scale_regularization=True,
max_gauss_ratio=10.0,
)
trainer = Trainer(config, output_dir=output_dir)
trainer.train() # ~10-40 minutes on GPU (A100: 10 min, T4: 25 min)
# Export to web-friendly format
export_gaussian_splat(output_dir / "splat.ply")
export_glb(output_dir / "model.glb") # for AR Quick Look / SceneViewer
return output_dir
Displaying the result in AR
// iOS: RealityKit Quick Look for .usdz / .glb
import RealityKit
import ARKit
class ModelViewerViewController: UIViewController {
func presentARModel(modelURL: URL) {
let arView = ARView(frame: view.bounds, cameraMode: .ar)
let anchor = AnchorEntity(plane: .horizontal)
ModelEntity.loadModelAsync(contentsOf: modelURL)
.sink(
receiveCompletion: { _ in },
receiveValue: { [weak self] entity in
entity.generateCollisionShapes(recursive: true)
anchor.addChild(entity)
arView.scene.anchors.append(anchor)
// Pinch to scale, pan to move
arView.installGestures([.scale, .translation, .rotation], for: entity)
}
)
.store(in: &cancellables)
}
}
Timeline and what's included in the work
| Stage | Duration | Result |
|---|---|---|
| Analysis and design | 3–5 days | Technical specification and architecture |
| Guided capture + validation | 5–7 days | Capture module with AR overlay |
| Cloud backend | 5–10 days | API for upload and training |
| AR viewing | 3–5 days | Integration of RealityKit / SceneViewer |
| Testing and deployment | 3–5 days | App in App Store / Google Play |
Note: what's included: full documentation, source code, cloud service deployment instructions, 3 months of support.
Why trust us with development?
We are certified iOS and Android specialists. With 5+ years of experience, we have delivered 30+ projects with computer vision and AR. We guarantee compliance with App Store Review Guidelines (Section 4.2/5.1) and stable operation with push notifications, deep linking, and in-app purchase. We provide a complete cycle: from idea to publication. Get an engineer consultation — we will help you choose the optimal pipeline for your task.
Contact us — we will evaluate your project within a day. We will answer your questions and offer the best turnkey solution.







