Furniture Search by Photo: AI Recognition in Mobile Apps

TRUETECH is engaged in the development, support and maintenance of iOS, Android, PWA mobile applications. We have extensive experience and expertise in publishing mobile applications in popular markets like Google Play, App Store, Amazon, AppGallery and others.

Development and support of all types of mobile applications:

Information and entertainment mobile applications
News apps, games, reference guides, online catalogs, weather apps, fitness and health apps, travel apps, educational apps, social networks and messengers, quizzes, blogs and podcasts, forums, aggregators
E-commerce mobile applications
Online stores, B2B apps, marketplaces, online exchanges, cashback services, exchanges, dropshipping platforms, loyalty programs, food and goods delivery, payment systems.
Business process management mobile applications
CRM systems, ERP systems, project management, sales team tools, financial management, production management, logistics and delivery management, HR management, data monitoring systems
Electronic services mobile applications
Classified ads platforms, online schools, online cinemas, electronic service platforms, cashback platforms, video hosting, thematic portals, online booking and scheduling platforms, online trading platforms

These are just some of the types of mobile applications we work with, and each of them may have its own specific features and functionality, tailored to the specific needs and goals of the client.

Showing 1 of 1All 1734 services
Furniture Search by Photo: AI Recognition in Mobile Apps
Complex
~1-2 weeks
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_mobile-applications_feedme_467_0.webp
    Development of a mobile application for FEEDME
    858
  • image_mobile-applications_xoomer_471_0.webp
    Development of a mobile application for XOOMER
    745
  • image_mobile-applications_rhl_428_0.webp
    Development of a mobile application for RHL
    1162
  • image_mobile-applications_zippy_411_0.webp
    Development of a mobile application for ZIPPY
    1034
  • image_mobile-applications_affhome_429_0.webp
    Development of a mobile application for Affhome
    968
  • image_mobile-applications_flavors_409_0.webp
    Development of a mobile application for the FLAVORS company
    563

Furniture Search by Photo: AI Recognition in Mobile Apps

A client wants users to snap a sofa and get similar models with prices. But furniture isn't clothing: visually identical chairs can differ in style and dimensions. A mistake in image classification leads to false recommendations. We solve this with a combination of transfer learning on convolutional networks and vector search. Our experience shows that a properly tuned model saves up to 40% of catalog manual processing time, and catalog maintenance costs drop by 25% through automation.

Why Furniture Is Harder Than Clothing

Furniture has rigid geometry and recognizable shapes — that simplifies classification. But searching for "similar for less money" requires not just visual match but also understanding style (Scandinavian minimalism, loft, classic) and scale. A sofa photo without context doesn't reveal whether it's a three-seater or a two-seater. We address this with additional furniture style classifiers and attribute filtering.

Source: recommendations from open furniture image datasets (e.g., Furniture-180).

Problems We Solve

  1. Categorization: Most off-the-shelf models (Google Cloud Vision, AWS Rekognition) are trained on ImageNet, which includes plenty of furniture. But accuracy drops for rare forms. We fine-tune EfficientNet or MobileNetV3 on the store's catalog: 50,000 images (200–300 per category) yield reliable classification of main categories: sofa, armchair, table, chair, wardrobe, bed, nightstand.
  2. Style detection: Category is just the first step. For similarity search, style matters. We use CLIP, which understands text descriptions: "Scandinavian minimalist", "loft", "classic". CLIP compares the image embedding with text style embeddings and returns the most likely.
  3. Filtered search: The architecture is the same as for clothing: embedding → vector search. But for furniture, filtering by material (wood, metal, fabric), color, and size group is crucial. Size class without AR is approximated by aspect ratio and typical category proportions.

Comparison of Classification Approaches

Approach Accuracy Training Time Dataset Size
Ready API (Google Cloud Vision) 75-85% 0 Not required
Transfer Learning (MobileNetV3) 92-96% 2-4 days 500–1000 per category
Training from scratch (ResNet) 96%+ 2-4 weeks 10,000+ per category

Transfer learning is the optimal choice for most retailers: 92% accuracy with minimal time investment. It outperforms ready APIs by 15% on average in our tests. And it requires 2-3 times less data than training from scratch at comparable accuracy. We use MobileNetV3 or EfficientNet-Lite — they run directly on the device.

How to Determine Furniture Style?

Method Style Accuracy Labeling Required
Visual features only (CNN) 70-80% 500+ images per style
CLIP (text+image) 85-95% Text descriptions of styles
Hybrid (CLIP + custom classifier) 90-97% 200 images per style

CLIP allows quick adaptation to a new style without retraining — just add a text description. More about CLIP in the official repository.

How We Do It: Stack and Examples

For Android we use TFLite with XNNPACK optimization. Example classifier:

// Android: TFLite furniture classifier
class FurnitureClassifier(context: Context) {

    private val interpreter: Interpreter by lazy {
        val model = FileUtil.loadMappedFile(context, "furniture_classifier_v2.tflite")
        Interpreter(model, Interpreter.Options().apply {
            numThreads = 4
            useXNNPACK = true
        })
    }

    fun classify(bitmap: Bitmap): List<FurnitureClassification> {
        val resized = Bitmap.createScaledBitmap(bitmap, 224, 224, true)
        val input = TensorImage.fromBitmap(resized)
        val output = TensorBuffer.createFixedSize(intArrayOf(1, NUM_CLASSES), DataType.FLOAT32)

        interpreter.run(input.buffer, output.buffer)

        return output.floatArray
            .mapIndexed { index, score -> FurnitureClassification(LABELS[index], score) }
            .filter { it.score > 0.1f }
            .sortedByDescending { it.score }
    }
}

For iOS — Core ML. CLIP enables style detection via text prompts:

// iOS: CLIP-based style detection via Core ML
// CLIP model converted to .mlpackage
func detectStyle(_ image: UIImage) async throws -> [StyleScore] {
    let styleDescriptions = [
        "scandinavian minimalist furniture",
        "industrial loft style furniture",
        "classic traditional furniture",
        "mid-century modern furniture",
        "boho eclectic furniture"
    ]

    // CLIP compares image embedding with text embeddings of styles
    let imageEmbedding = try await clipEncoder.encodeImage(image)
    return styleDescriptions.enumerated().map { i, desc in
        let textEmbedding = clipEncoder.encodeText(desc)
        let similarity = cosineSimilarity(imageEmbedding, textEmbedding)
        return StyleScore(style: desc, score: similarity)
    }.sorted { $0.score > $1.score }
}

How Does Similar Item Search Work in the Catalog?

Product embeddings are stored in a vector database (e.g., FAISS or Pinecone). Search involves several steps:

  1. Extract embedding: the user's image goes through the same model as the catalog.
  2. Vector search: cosine distance to all catalog embeddings.
  3. Apply filters: filter by category, style, material, color, size, and price.
  4. Post-processing: sort by similarity and return top-10 results.
struct FurnitureSearchFilters {
    let category: FurnitureCategory
    let style: StyleTag?
    let colorFamily: ColorFamily?        // warm, cool, neutral
    let material: MaterialType?          // wood, metal, upholstered
    let maxDimensionClass: SizeClass?    // compact, standard, large
    let priceRange: ClosedRange<Int>?
    let inStockOnly: Bool
}

Integration Process

Implementing recognition and similar-item search goes through several stages:

More details on stages
  1. Catalog analysis: estimate image volume, categories, styles.
  2. Data preparation: label 200–500 images per category/style (if fine-tuning is needed).
  3. Model training: transfer learning on TensorFlow/Keras or Create ML.
  4. App integration: connect TFLite/Core ML, set up vector storage.
  5. Testing: A/B test recognition accuracy and search speed.
  6. Deployment: release via App Store/Google Play, monitor.

A minimum viable product (MVP) can be launched in 1 week using a ready-made recognition API and cloud vector storage. A full solution with a custom model, CLIP, and AR measures takes 1–2 months. For example, for a client with 10,000 catalog items, we fine-tuned MobileNetV3 on 500 images per category. Classification accuracy rose from 72% to 94%.

Scope of Work

  • MVP: integration with a ready API (Google Cloud Vision + vector search) — from 1 week.
  • Full solution: fine-tuned TFLite/CoreML model, CLIP style classifier, vector store with filters, AR dimensions (LiDAR) — 1–2 months.
  • Model and API documentation.
  • Client team training.
  • Post-release support for 2 weeks.

Our team has 7+ years of experience in mobile development and machine learning. We have completed 30+ projects in image recognition. We guarantee quality: the model will work offline without delays. If you want to implement a similar feature, contact us — we'll find the optimal solution. Get a free consultation on integration — we'll assess your project at no cost.

Machine Learning in Mobile Apps: CoreML, TFLite, and On-Device Models

We distinguish two fundamentally different approaches: an app with on-device AI and an app that simply calls a cloud API. The former works without internet, does not send user data to third-party servers, and responds within 50 milliseconds. The latter depends on network latency and pricing plans. Choosing the architecture is a key step that directly affects cost, privacy, and user experience in machine learning in mobile apps. Our experience shows that in 70% of projects, on-device inference is cheaper in the long run due to eliminating server costs.

How to Choose Between CoreML and TFLite for On-Device Inference?

CoreML — Apple's native framework for running ML models on device. Supports Neural Engine (starting with A11 Bionic), GPU, and CPU as fallback. Models are converted to .mlmodel format via coremltools from PyTorch, ONNX, or TensorFlow. Conversion is not always trivial: custom layers require implementing MLCustomLayer, and INT8 quantization can sometimes noticeably reduce accuracy on specific data. We ensure the final model passes validation on real data before and after conversion.

TensorFlow Lite — cross-platform alternative for Android and Flutter. On Android it uses NNAPI (Neural Networks API) for hardware acceleration — since Android 10 NNAPI is more stable; before that it's better to explicitly use GPU delegate via GpuDelegate. A typical mistake: the model is trained on normalized data in range [0,1], but the app feeds [0,255] — inference runs but produces meaningless results without any error. We include an automatic input data validation module in the SDK.

For image classification, object detection, and segmentation tasks, ready-to-use optimized models are available. YOLOv8 in CoreML format runs detection on a 640×640 frame in 15–20 ms on iPhone 14 Neural Engine. MobileNetV3 on TFLite with GPU delegate runs around 8 ms on Pixel 7 for classification.

Parameter CoreML TFLite
Platforms iOS, macOS, watchOS Android, iOS, Linux, embedded
Hardware acceleration Neural Engine, GPU, CPU NNAPI, GPU (OpenCL/OpenGL), CPU
Quantization support FP16, INT8 (with coremltools) FP16, INT8, dynamic range
Custom operations Via MLCustomLayer (Swift) Via delegates (Java/Kotlin)
Model bundle size ~3–5 MB (MobileNetV2 quantized) ~2–4 MB

What If You Need Text Generation On-Device?

Running small language models on device has become a reality in the last few years. Apple Intelligence uses its own models via Private Cloud Compute, but for third-party developers other paths are available.

llama.cpp with Metal backend on iOS is a working approach for phi-3-mini (3.8B parameters, 4-bit quantization, ~2.3 GB). Inference: 15–25 tokens/second on iPhone 15 Pro. For integration in Swift, use the Swift Package llama.swift or a wrapper via C interface llama.h. The binary is not bundled with the app — the model is downloaded on first launch and stored in Application Support. Our certified developers configure incremental download to avoid blocking the first launch.

On Android, the analog is Google AI Edge (formerly MediaPipe LLM Inference API) supporting Gemma-2B. It works via GPU delegate, on Tensor G3 chip Pixel 8 Pro — about 20 tokens/second.

Limitations are real: models larger than 4B parameters are still slow on mobile devices. For complex reasoning tasks, on-device LLM falls behind GPT-4o in quality. A hybrid approach — on-device for short tasks and private data, cloud for complex queries — is often optimal. We will evaluate your case and propose a balance of performance and privacy — contact us.

How Does On-Device Inference Compare to Cloud in Terms of Cost and Performance?

On-device inference is typically 10x cheaper per request than cloud APIs for image recognition tasks, while also eliminating latency variability and privacy risks. The table below summarizes the trade-offs.

Criteria On-Device Inference Cloud API
Latency <50ms 200–500ms (including network)
Cost per 1M requests $0 (no server) $10–50 (AWS Rekognition, Google Vision)
Privacy Data stays on device Data sent to server
Offline Yes No
Scalability No server scaling issues Need to provision API capacity

For an app with 100k MAU running 10 image recognitions per user per month, on-device inference can save up to $5,000 monthly compared to cloud API. Get a free consultation on your ML architecture today.

Integrating OpenAI API and Other Cloud Models

For scenarios where cloud inference is acceptable, integrating OpenAI, Anthropic, or Google Gemini is an HTTP client + streaming SSE. In Swift, AsyncThrowingStream is convenient for streaming responses. In Kotlin, use Flow.

Critically: API keys must never be stored in the app bundle. Even an obfuscated key can be extracted from the IPA in 10 minutes using strings or frida. Correct architecture: mobile app → your own backend → OpenAI API. The backend controls rate limiting, logs requests, and protects the key.

What Is Included in the Work (Deliverables)

  • Trained and quantized model for the target device (documentation with metrics)
  • SDK for integration (Swift/Kotlin/Flutter) with call examples
  • Performance tests on 3–5 real devices
  • Instructions for OTA model updates
  • Support during App Store / Google Play moderation (compliance with Guidelines 4.2, 5.1)
  • 2 weeks of technical support after release

Typical Project Pipeline

  1. Task analysis — measure latency, privacy, size, supported devices.
  2. Model prototyping — in Python, evaluate accuracy on target data.
  3. Conversion and quantization — for CoreML/TFLite with validation.
  4. Integration into the app — model wrapped in a service layer (easy to swap CoreML ↔ TFLite ↔ cloud).
  5. Testing — on real devices, measure FPS, RAM, battery.
  6. Deployment — via TestFlight / Firebase App Distribution, monitor metrics.

Timelines: integration of a ready CoreML/TFLite model — 1–2 weeks, development of a custom model with mobile optimization — from 6 weeks, on-device LLM chat with personalization — 4–8 weeks.

Why We Take on Complex Cases?

10+ years of experience in mobile development, 50+ implemented AI/ML solutions, guarantee of compatibility with current iOS and Android versions. All projects undergo code review and load testing. The cost includes preparation of moderation documentation and training of your team.

Contact us — we will help you choose the architecture and implement ML in your app turnkey. Order an audit of your existing solution — we will assess the potential for server cost savings free of charge. In some projects, savings can reach significant amounts per month.