Automatic Camera-Based Inventory System
Imagine a warehouse with 10,000 SKUs—manual inventory every quarter halts shipments for three days. Discrepancies are discovered after the fact, leaving no time to correct them. We solve this with a computer vision system: cameras analyze shelves in real time, and YOLO algorithms (Ultralytics YOLOv8) detect each item. The system runs without stopping warehouse operations and achieves up to 98% accuracy, which is 2 times better than typical manual accuracy of 95%. For a 10,000-SKU warehouse, savings exceed 2 million rubles per year (approx. $24,000). This translates to annual savings of $24,000 in inventory labor costs. Our company has 5+ years of experience in warehouse automation and has completed 50+ projects, guaranteeing at least 95% accuracy from the first run.
How Camera Inventory Works
Fixed cameras above shelves take snapshots on a schedule (every 6–12 hours). Images pass through an object detection model—we use YOLOv8 (trained on 100+ product classes). The output is a list of SKUs with counts per shelf. These are reconciled with expected stock levels from your ERP. Discrepancies are flagged automatically, and the system can trigger replenishment orders. YOLOv8 processes a frame twice as fast as its predecessor, critical for hundreds of shelves. Optionally, we use NVIDIA Triton for batch inference, reducing p99 latency to 50ms.
Why 92–96% Accuracy Isn't Enough and How We Boost It
For most retailers, 95% accuracy seems acceptable, but at million-dollar turnover, each percentage point of discrepancy means direct losses. We push accuracy to 98% with three techniques: multi-angle capture, data augmentation, and reconciliation with POS data. Multi-angle capture reduces occlusion errors by 40%.
Achieving 98% Accuracy
Multi-angle capture is key. One camera cannot see items hidden behind others, so we install 2–3 cameras per aisle. For each frame, we apply perspective transformation to get a flat shelf view. Then we merge detections from different angles, removing duplicates by IoU (Intersection over Union). This reduces occlusion errors by 40%.
Reconciliation with the Accounting System
def reconcile(camera_counts: dict, system_counts: dict,
tolerance_percent: float = 5.0) -> list[dict]:
"""Find discrepancies between physical and system counts"""
discrepancies = []
all_skus = set(camera_counts) | set(system_counts)
for sku in all_skus:
camera_qty = camera_counts.get(sku, 0)
system_qty = system_counts.get(sku, 0)
if system_qty > 0:
diff_pct = abs(camera_qty - system_qty) / system_qty * 100
else:
diff_pct = 100 if camera_qty > 0 else 0
if diff_pct > tolerance_percent:
discrepancies.append({
'sku': sku,
'camera': camera_qty,
'system': system_qty,
'diff_percent': round(diff_pct, 1),
'severity': 'high' if diff_pct > 20 else 'medium'
})
return sorted(discrepancies, key=lambda x: x['diff_percent'], reverse=True)
Architecture & Tech Stack
The core of the system is a YOLO-based detector (PyTorch). Images are sent to a server with a GPU (NVIDIA T4 or A10); inference takes <100ms per image. We apply OpenCV perspective transformation to get a flat shelf view. For edge devices, we use TensorRT, speeding inference by 1.5× without accuracy loss.
class AutoInventorySystem:
def __init__(self, detector_path: str, inventory_db_path: str):
self.detector = YOLO(detector_path)
self.db = InventoryDatabase(inventory_db_path)
def run_inventory_cycle(self, shelf_images: dict) -> InventoryReport:
"""
shelf_images: {shelf_id: image} - photos of all shelves
"""
report = InventoryReport()
for shelf_id, image in shelf_images.items():
shelf_counts = self._count_shelf(image, shelf_id)
report.add_shelf(shelf_id, shelf_counts)
# Compare with expected stock
expected = self.db.get_expected_quantities()
report.discrepancies = self._find_discrepancies(
report.actual_counts, expected
)
# Automatic update in ERP
self.db.update_inventory(report.actual_counts)
return report
def _count_shelf(self, image: np.ndarray,
shelf_id: str) -> dict:
"""Count products on one shelf"""
detections = self.detector(image, conf=0.45)
counts = {}
for box in detections[0].boxes:
sku = self.detector.model.names[int(box.cls)]
counts[sku] = counts.get(sku, 0) + 1
return counts
Perspective Distortion Handling
def create_shelf_rectified_view(image: np.ndarray,
shelf_corners: list,
output_size: tuple = (2000, 400)) -> np.ndarray:
"""
Flat (top-down) representation of the shelf for easier analysis
shelf_corners: 4 corners of the shelf in the image
"""
pts_src = np.array(shelf_corners, dtype='float32')
w, h = output_size
pts_dst = np.array([
[0, 0], [w - 1, 0],
[w - 1, h - 1], [0, h - 1]
], dtype='float32')
M = cv2.getPerspectiveTransform(pts_src, pts_dst)
rectified = cv2.warpPerspective(image, M, (w, h))
return rectified
Drone-Based Inventory
class DroneInventoryController:
def __init__(self, drone_api, inventory_system):
self.drone = drone_api
self.inventory = inventory_system
self.waypoints = [] # pre-programmed shooting points
async def run_inventory_mission(self) -> InventoryReport:
images = {}
await self.drone.takeoff()
for waypoint in self.waypoints:
await self.drone.fly_to(waypoint)
await self.drone.stabilize(seconds=1.0)
# Photos from multiple angles for better coverage
for angle_offset in [0, -15, 15]:
await self.drone.rotate(angle_offset)
image = await self.drone.capture_image()
images[f"{waypoint['shelf_id']}_{angle_offset}"] = image
await self.drone.land()
return self.inventory.run_inventory_cycle(images)
ERP/WMS Integration
import requests
class ERPIntegration:
def __init__(self, erp_url: str, api_key: str):
self.base_url = erp_url
self.headers = {'Authorization': f'Bearer {api_key}'}
def update_stock_levels(self, inventory: dict,
location_id: str) -> dict:
"""Update stock levels in the ERP system"""
stock_updates = [
{
'sku': sku,
'quantity': qty,
'location_id': location_id,
'source': 'camera_inventory',
'timestamp': get_iso_timestamp()
}
for sku, qty in inventory.items()
]
response = requests.post(
f'{self.base_url}/api/inventory/bulk-update',
json={'updates': stock_updates},
headers=self.headers
)
return response.json()
| Inventory Type | Accuracy | Time |
|---|---|---|
| Fixed cameras (retail) | 92–96% | Continuous |
| Drone (1000 m² warehouse) | 90–95% | 20–40 min |
| Mobile robot | 94–98% | 30–60 min |
Implementation Process
- Analysis: we survey your warehouse, determine shelf count, angles, lighting. Select equipment.
- Design: we architect the system (cameras → server → ERP). Fine-tune the YOLO model for your product range.
- Implementation: mount cameras, deploy software, write integrations.
- Test: pilot run, reconcile with manual inventory, calibrate.
- Deploy: go live, train staff.
- Monitor & optimize: after launch, track accuracy, retrain the model for new SKUs as needed.
Deliverables
- Documentation: camera installation diagram, API spec, operator manual.
- Training: 2 days for warehouse team and 1 day for IT staff.
- Support: 12-month warranty on software, 3 months of free support.
- Metrics: dashboard showing accuracy, discrepancy count, inventory time.
Typical Mistakes and How We Avoid Them
- Occlusion: products hide each other. Solution: multi-angle capture.
- Changing lighting: we retrain the model on night-time frames.
- New SKUs: automatic registration via image search (embeddings).
| Scale | Timeline |
|---|---|
| Warehouse/store with fixed cameras | 6–9 weeks |
| Drone system + ERP | 10–16 weeks |
| Full autonomous system | 16–24 weeks |
Contact us for a free preliminary assessment of your warehouse. Get an individual estimate and pilot project—see the accuracy on real data. With our experience (5+ years, 50+ projects), we guarantee at least 95% accuracy from the first run. Time savings on inventory: up to 80%; loss reduction from discrepancies: up to 30%. Typical project cost starts from $50,000 for a warehouse with 10,000 SKUs.







