Browser Automation with AI: Computer Vision & LLM

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Browser Automation with AI: Computer Vision & LLM
Medium
~2-4 weeks
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Browser Automation with AI: Computer Vision & LLM

Imagine: a competitor's site updates its interface, and your RPA bot, tied to selectors like #price-block, starts throwing errors. Sound familiar? We solve this with AI agents that analyze the interface like a human — via screenshots and element semantics. They don't break on frontend refactoring and are resilient to anti-bot protection. Over 5+ years, we've delivered 30+ automation projects for e-commerce, logistics, and finance, saving clients an average of 70% manual work.

Traditional RPA bots are rigidly tied to DOM structure. Change a button's class — the script crashes. AI agents analyze screenshots and element semantics, so they work with any interface — SPA, React, Vue, dynamic elements. Flexibility is orders of magnitude higher: frontend refactoring doesn't break the agent.

Characteristic RPA Bot AI Agent
Resistance to layout changes Low (breaks on class changes) High (semantic search)
CAPTCHA handling Requires anti-captcha integration Can solve via visual analysis
Setup complexity High (precise selectors) Low (task in natural language)
Scalability Only predefined scenarios Adapts to new pages

What is an AI Agent and How is It Different from RPA?

An AI agent is a software bot that uses a multimodal large language model (LLM) and computer vision to interact with web interfaces. Unlike RPA, which follows rigid selectors, an AI agent sees the page like a human: it recognizes buttons, input fields, and other elements by their visual appearance and context. This makes it resilient to layout changes, framework swaps, and even some anti-bot protections.

How We Build the Agent: Architecture

We use a combination of Playwright + LLM (Claude, GPT‑4). Playwright controls the browser; the LLM makes decisions — selecting the next action based on the screenshot and available elements.

from playwright.async_api import async_playwright, Page
from anthropic import Anthropic
import asyncio
import base64
import json

client = Anthropic()

class AIWebAgent:

    def __init__(self, headless: bool = True):
        self.headless = headless

    async def run(self, task: str, start_url: str) -> str:
        async with async_playwright() as p:
            browser = await p.chromium.launch(headless=self.headless)
            context = await browser.new_context(
                viewport={"width": 1280, "height": 800},
                user_agent="Mozilla/5.0 (compatible; research-bot/1.0)"
            )
            page = await context.new_page()
            await page.goto(start_url)
            result = await self._agent_loop(page, task)
            await browser.close()
            return result

    async def _get_page_state(self, page: Page) -> dict:
        screenshot = await page.screenshot(type="png")
        screenshot_b64 = base64.standard_b64encode(screenshot).decode()

        elements = await page.evaluate("""() => {
            const interactive = document.querySelectorAll(
                'a, button, input, select, textarea, [role="button"]'
            );
            return Array.from(interactive).slice(0, 50).map((el, idx) => ({
                index: idx,
                tag: el.tagName.toLowerCase(),
                text: el.textContent?.trim().slice(0, 80) || '',
                type: el.type || '',
                placeholder: el.placeholder || '',
            }));
        }""")

        return {
            "url": page.url,
            "title": await page.title(),
            "screenshot": screenshot_b64,
            "elements": elements,
        }

    async def _execute_action(self, page: Page, action: dict) -> str:
        action_type = action.get("type")
        try:
            if action_type == "click":
                if "text" in action:
                    await page.get_by_text(action["text"]).first.click()
                elif "element_index" in action:
                    elements = await page.query_selector_all(
                        'a, button, input, select, textarea, [role="button"]'
                    )
                    if action["element_index"] < len(elements):
                        await elements[action["element_index"]].click()
            elif action_type == "fill":
                await page.fill(action["selector"], action["value"])
            elif action_type == "navigate":
                await page.goto(action["url"])
            elif action_type == "scroll":
                amount = action.get("amount", 500)
                direction = 1 if action.get("direction", "down") == "down" else -1
                await page.evaluate(f"window.scrollBy(0, {direction * amount})")
            elif action_type == "extract":
                return await page.evaluate(
                    f'document.querySelector({json.dumps(action.get("selector", "body"))})?.textContent?.trim()'
                )
            await asyncio.sleep(0.5)
            return "success"
        except Exception as e:
            return f"error: {e}"

    async def _agent_loop(self, page: Page, task: str) -> str:
        messages = []
        system = """You are an AI web agent. Complete tasks in the browser.

Available actions (return JSON):
- {"type": "click", "text": "button text"}
- {"type": "fill", "selector": "CSS", "value": "text"}
- {"type": "navigate", "url": "https://..."}
- {"type": "scroll", "direction": "down", "amount": 500}
- {"type": "extract", "selector": ".result"}
- {"type": "done", "result": "final result"}

Response format:
REASONING: <what you see and plan>
ACTION: <JSON with one action>"""

        for _ in range(20):
            state = await self._get_page_state(page)
            user_content = [
                {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": state["screenshot"]}},
                {"type": "text", "text": f"Task: {task}\nURL: {state['url']}\nElements:\n{json.dumps(state['elements'][:20], ensure_ascii=False)}"}
            ]
            messages.append({"role": "user", "content": user_content})

            response = client.messages.create(
                model="claude-opus-4-5",
                max_tokens=1024,
                system=system,
                messages=messages,
            )
            reply = response.content[0].text
            messages.append({"role": "assistant", "content": reply})

            try:
                action_line = [l for l in reply.split("\n") if l.startswith("ACTION:")][0]
                action = json.loads(action_line.replace("ACTION:", "").strip())
            except (IndexError, json.JSONDecodeError):
                continue

            if action.get("type") == "done":
                return action.get("result", "Task completed")

            await self._execute_action(page, action)

        return "Max iterations reached"
Technical implementation of anti-bot protection
async def setup_stealth_context(playwright):
    browser = await playwright.chromium.launch(
        headless=False,
        args=["--disable-blink-features=AutomationControlled"]
    )
    context = await browser.new_context(
        viewport={"width": 1366, "height": 768},
        locale="en-US",
        timezone_id="America/New_York",
    )
    await context.add_init_script(
        "Object.defineProperty(navigator, 'webdriver', {get: () => undefined})"
    )
    return context

Modern websites actively block automation. We add navigator.webdriver modification, a realistic User-Agent, random delays between actions, and proxy rotation. The agent can even bypass sophisticated systems — Cloudflare, DataDome, Akamai.

Why an AI Agent is Faster Than a Human

A person spends about 4 hours per day monitoring 200 products from 5 competitors. An AI agent does the same in 35 minutes and produces an Excel report. The response time to price changes shrinks from 2 days to 2 hours.

How the AI Agent Handles CAPTCHA

The agent takes a screenshot of the CAPTCHA element and sends it to a multimodal LLM. The model recognizes the text or selects the correct images. This eliminates the need for third-party anti-captcha services and speeds up solving — on average 3–5 seconds per check.

Typical Use Cases

  • Competitor monitoring — prices, promotions, assortment
  • Form filling — tenders, registrations, government services
  • Review aggregation — collecting feedback from multiple platforms
  • UI regression testing — automatic scenario verification

Case Study: Building Materials Retail Chain

Our client — a building materials retail chain — manually monitored prices for 200 SKUs from 5 competitors every day. We deployed an AI agent: it browses websites, extracts prices and availability, and generates an Excel file with deviations. The task runs in 35 minutes. Result: the manager spends 15 minutes analyzing the ready report instead of 4 hours collecting data. Price change reaction time improved from 1–2 days to a few hours.

What's Included in the Work

  • Development of an AI agent for your specific scenario
  • Source code and documentation in English
  • Anti-bot protection and proxy setup (if needed)
  • Integration with your systems via API or data export
  • Employee training on using the agent
  • Technical support for the first month

Work Process

  1. Analysis — we study target sites, anti-bot protection, data structure
  2. Design — we choose the stack (model, framework, proxy)
  3. Implementation — we code the agent and set up the pipeline
  4. Testing — we run through a set of complex scenarios
  5. Deployment — we place it on a server with monitoring

Estimated Timelines

  • Basic agent for one site: 3–5 days
  • Universal agent with visual perception: 1–2 weeks
  • Price monitoring with reports: 1 week
  • Anti-bot protection and proxies: +1 week

Cost starts at $2,500 per basic agent, with average savings of $15,000/year in manual work. We guarantee stable operation as long as target sites remain unchanged. We'll assess your project in 1 day — contact us. Order AI agent development for your business. Get a consultation on your automation scenario — write to us.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.