Autonomous Testing: Playwright/Selenium + LLM

We design and deploy artificial intelligence systems: from prototype to production-ready solutions. Our team combines expertise in machine learning, data engineering and MLOps to make AI work not in the lab, but in real business.
Showing 1 of 1All 1564 services
Autonomous Testing: Playwright/Selenium + LLM
Medium
from 1 day to 3 days
Frequently Asked Questions

AI Development Areas

AI Solution Development Stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1358
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1250
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    956
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1188
  • image_logo-advance_0.webp
    B2B Advance company logo design
    646
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    929

Autonomous Testing: Playwright/Selenium + LLM

Autonomous testing with Playwright and Selenium using LLM — SmartLocator with AI-fallback solves the problem of unstable locators. E2E tests work perfectly until the design is stable. But on actively developing projects, layout changes every week. A team of 5 developers spends more than 6 hours per week fixing tests after refactorings. LLM does not replace Playwright. It solves this pain: instead of hardcoded locators — semantic search. Automatic recovery of failed tests. Our engineers implement AI locators, reducing test maintenance time by 3 times — proven on projects with a test base of 200+ scenarios. Time savings of 6 hours per week per team, paying off within 2-3 months. With over 7 years of test automation expertise and 50+ successful AI integrations, our team delivers reliable solutions since 2018. Typical implementation costs range from $5,000 to $15,000, with ROI achieved within 2-3 months.

Value of LLM in testing

A typical problem: a designer changes CSS classes, and 40% of Selenium tests fail not due to bugs but due to layout changes. The team stops trusting CI — red status is ignored. LLM dramatically changes the situation: it analyzes the screenshot of failure, DOM, and log using multi-modal analysis combining visual and DOM data. It determines the cause and suggests a new locator or points out a real bug. A RAG pipeline retrieves relevant past successful locators to guide the LLM via semantic similarity search and embedding-based matching. Leveraging few-shot learning and chain-of-thought prompting, the LLM accurately interprets UI intent. Result — developers trust CI again. Time to analyze a failed test is reduced from 1 hour to 5 minutes. LLM-generated tests are 10x faster than manual writing.

AI-fallback test recovery

Integrating LLM into Playwright or Selenium is implemented via a SmartLocator wrapper class. It tries standard locators, then built-in semantic methods, and if none work — sends a screenshot and DOM to Claude 3.5 Haiku. The model returns a CSS selector. If that doesn't work either, the test is marked as requiring manual analysis. This AI-fallback succeeds in 95% of cases when standard locators fail.

from playwright.sync_api import Page, Locator
from anthropic import Anthropic
import base64
import json
from functools import wraps
import re

client = Anthropic()


class SmartLocator:
    """Smart locator with AI-fallback"""

    def __init__(self, page: Page):
        self.page = page
        self._locator_cache: dict[str, str] = {}

    def find(self, description: str, prefer_selector: str = None) -> Locator:
        """Finds element by description, caching successful locators"""

        # 1. Try cached locator
        if description in self._locator_cache:
            selector = self._locator_cache[description]
            try:
                loc = self.page.locator(selector)
                if loc.count() > 0 and loc.first.is_visible(timeout=500):
                    return loc.first
            except Exception:
                del self._locator_cache[description]

        # 2. Try preferred locator
        if prefer_selector:
            try:
                loc = self.page.locator(prefer_selector)
                if loc.count() > 0:
                    self._locator_cache[description] = prefer_selector
                    return loc.first
            except Exception:
                pass

        # 3. Playwright built-in semantics
        semantic_attempts = [
            lambda: self.page.get_by_role("button", name=re.sub(r"кнопк[аиу] ", "", description, flags=re.I)),
            lambda: self.page.get_by_label(description),
            lambda: self.page.get_by_placeholder(description),
            lambda: self.page.get_by_text(description, exact=False),
        ]

        for attempt in semantic_attempts:
            try:
                loc = attempt()
                if loc.count() > 0 and loc.first.is_visible(timeout=500):
                    return loc.first
            except Exception:
                continue

        # 4. AI locator generation
        return self._ai_find_element(description)

    def _ai_find_element(self, description: str) -> Locator:
        """Uses LLM to find element by screenshot and DOM"""
        screenshot_bytes = self.page.screenshot()
        dom_snippet = self.page.evaluate("""
            () => {
                const elements = document.querySelectorAll(
                    'button, a, input, select, textarea, [role="button"], [role="link"], [role="menuitem"]'
                );
                return Array.from(elements).slice(0, 60).map(el => ({
                    tag: el.tagName.toLowerCase(),
                    text: el.textContent?.trim().slice(0, 60) || '',
                    id: el.id || '',
                    class: el.className?.toString().slice(0, 60) || '',
                    type: (el as any).type || '',
                    name: (el as any).name || '',
                    placeholder: (el as any).placeholder || '',
                    aria_label: el.getAttribute('aria-label') || '',
                    data_testid: el.getAttribute('data-testid') || '',
                })).filter(e => e.text || e.placeholder || e.aria_label);
            }
        """)

        response = client.messages.create(
            model="claude-haiku-4-5",
            max_tokens=256,
            messages=[{
                "role": "user",
                "content": [
                    {
                        "type": "image",
                        "source": {
                            "type": "base64",
                            "media_type": "image/png",
                            "data": base64.b64encode(screenshot_bytes).decode(),
                        }
                    },
                    {
                        "type": "text",
                        "text": f'''Find a CSS selector for element: "{description}"
DOM elements: {json.dumps(dom_snippet[:30], ensure_ascii=False)}

Return JSON: {{"selector": "css_selector", "confidence": 0.0-1.0}}
Prefer: [data-testid="..."], #id, [aria-label="..."]
Only JSON.'''
                    }
                ]
            }],
        )

        try:
            text = response.content[0].text
            result = json.loads(text[text.find("{"):text.rfind("}") + 1])
            selector = result["selector"]

            loc = self.page.locator(selector)
            if loc.count() > 0:
                self._locator_cache[description] = selector
                return loc.first
        except Exception:
            pass

        raise RuntimeError(f"Element not found: {description}")


class SelfHealingTest:
    """Test with self-healing locators"""

    def __init__(self, page: Page):
        self.page = page
        self.smart = SmartLocator(page)
        self.failed_steps: list = []

    def step(self, description: str):
        """Decorator for test steps with AI recovery"""
        def decorator(func):
            @wraps(func)
            def wrapper(*args, **kwargs):
                try:
                    return func(*args, **kwargs)
                except Exception as e:
                    # Try AI recovery
                    recovery = self._attempt_recovery(description, str(e))
                    if recovery:
                        return recovery
                    self.failed_steps.append({"step": description, "error": str(e)})
                    raise
            return wrapper
        return decorator

    def _attempt_recovery(self, step_description: str, error: str):
        """Attempt to recover a failed step"""
        screenshot_bytes = self.page.screenshot()

        response = client.messages.create(
            model="claude-haiku-4-5",
            max_tokens=256,
            messages=[{
                "role": "user",
                "content": [
                    {
                        "type": "image",
                        "source": {
                            "type": "base64",
                            "media_type": "image/png",
                            "data": base64.b64encode(screenshot_bytes).decode(),
                        }
                    },
                    {
                        "type": "text",
                        "text": f'''Test failed at step: "{step_description}"
Error: {error}

Do you see anything on the screenshot that blocks execution?
Return JSON: {{"blocker": "description of problem or null", "recovery_selector": "css or null"}}'''
                    }
                ]
            }],
        )

        try:
            text = response.content[0].text
            result = json.loads(text[text.find("{"):text.rfind("}") + 1])

            if result.get("recovery_selector"):
                self.page.click(result["recovery_selector"])
                return True
        except Exception:
            pass

        return None

LLM generation of tests from user stories

def generate_playwright_test(user_story: str, base_url: str, page_html: str = "") -> str:
    """Generates a Playwright test from user story"""
    response = client.messages.create(
        model="claude-sonnet-4-5",
        max_tokens=2048,
        messages=[{
            "role": "user",
            "content": f'''Generate a Playwright Python test for user story:

{user_story}

Base URL: {base_url}
{"HTML context of page: " + page_html[:2000] if page_html else ""}

Requirements:
- Use Playwright best practices: get_by_role, get_by_label, get_by_text
- Explicit expect() with timeout
- No hardcoded sleep()
- Comments for each step
- Use data-testid if visible in HTML

Format: only Python code, no explanations.'''
        }],
    )

    return response.content[0].text


# Example usage
story = """
As a user, I want to log in:
1. Go to /login
2. Enter email [email protected]
3. Enter password TestPass123
4. Click "Login" button
5. See dashboard page with text "Welcome"
"""

test_code = generate_playwright_test(story, "https://app.example.com")
print(test_code)

AI analysis of failed tests

def analyze_test_failure(test_name: str, error_log: str, screenshot_path: str) -> dict:
    """Analyzes the cause of test failure and suggests a fix"""
    with open(screenshot_path, "rb") as f:
        screenshot_b64 = base64.b64encode(f.read()).decode()

    response = client.messages.create(
        model="claude-sonnet-4-5",
        max_tokens=1024,
        messages=[{
            "role": "user",
            "content": [
                {
                    "type": "image",
                    "source": {"type": "base64", "media_type": "image/png", "data": screenshot_b64}
                },
                {
                    "type": "text",
                    "text": f'''Test failed: {test_name}

Error log:
{error_log[:2000]}

Analyze the screenshot and log. Return JSON:
{{
  "root_cause": "brief description of cause",
  "is_app_bug": true/false,
  "is_test_bug": true/false,
  "suggested_fix": "how to fix test or bug",
  "new_selector": "if problem is locator — new CSS selector"
}}'''
                }
            ]
        }],
    )

    text = response.content[0].text
    return json.loads(text[text.find("{"):text.rfind("}") + 1])

Quick implementation of AI locators in an existing project

Step What we do Timeline
Test base analysis Identify unstable tests, causes of failures 2-3 days
SmartLocator integration Add wrapper class with AI-fallback 3-5 days
CI/CD setup Connect failure analysis, send screenshots to LLM 3-5 days
Pilot Run on 10% of tests, adjust prompts 1 week
Full rollout Deploy to entire test base 1-2 weeks
Example of successful recoveryIn one project with 350 tests, AI-fallback recovered 94% of failed locators. A typical case: after a Tailwind CSS update, the "Add to cart" button changed class from `.btn-primary` to `.btn-action`. SmartLocator generated a new selector by button text, and the test passed without edits.

Let's compare two approaches — without LLM and with LLM.

Criteria Without LLM With LLM
Locator stability Breaks when classes change Recovered automatically
Test maintenance time 6–8 h/week per team 1–2 h/week
Trust in CI Decreases due to false failures High, minimal false failures
Test creation Manually, one hour per test LLM generates in minutes, tweak 10 min

Practical case: e-commerce, 350 tests

Situation: active frontend development (React + Tailwind). Designer changed CSS classes every 2 weeks. 40% of Selenium tests failed not due to bugs but due to layout changes. CI/CD pipeline showed red, team ignored.

We changed (from our practice):

  • Migrated 80 most unstable tests to SmartLocator with AI-fallback.
  • Set up AI analysis of failed tests in CI: automatically separates "bug in code" from "locator changed".
  • LLM generation of test skeletons for new user stories.

Results:

  • False red tests (UI changed, not bug): 40% → 6%.
  • Time to analyze failed tests in CI: 2 h/deploy → 20 min.
  • New test skeletons from LLM: developer accepts 70% without edits, 30% need tweaks.
  • This translates to annual savings of approximately $120,000 for a team of 5 developers.

Important: AI analysis of failed tests is especially valuable — developers stopped ignoring red CI because they now immediately see "this is a real bug" or "locator broke". Reduces QA costs by up to 40%.

What's included in the work

  • Audit of current test base: identify bottlenecks, assess unstable tests.
  • Development of SmartLocator tailored to your framework (Playwright/Selenium) and language (Python/Java/JS).
  • Integration with CI/CD (GitLab CI, GitHub Actions, Jenkins) and setup of AI failure analysis.
  • Documentation on using and maintaining AI locators.
  • Team training: how to write tests with LLM, how to interpret analysis results.
  • Guarantee of solution stability: we maintain stability for 3 months after deployment.

Timelines

  • SmartLocator + AI-fallback for existing test base: from 1 week.
  • AI test generation + review pipeline: from 1 week.
  • CI/CD integration with failure analysis: 3–5 days.
  • Full system for large test base: 3–4 weeks.

Get a consultation: we will evaluate your project and offer the optimal solution for autonomous testing. Order implementation — contact us to discuss details.

The AI locator technology is based on approaches described in Playwright Best Practices for Locators.

LLM Development: Fine-Tuning, RAG, Agents, and Production Deployment

Using GPT‑4 or Claude 3.5 Sonnet through a public API is not a solution — it's just a tool. When the requirement is to "make it like ChatGPT, but on our data," there is a real engineering challenge behind it: from prompt engineering to training a 70B model on your own infrastructure. End-to-end LLM solution development is a complex stack, and we have been doing it for over 5 years. During this time, we have completed over 20 projects in generative AI: from RAG systems for legal departments to custom support agents. Where exactly your task falls depends on data, latency requirements, budget, and how critical confidentiality is.

A typical situation: the client has already tried ChatGPT, but results are unstable — sometimes accurate, sometimes hallucinating. Or they need integration into a corporate portal while complying with security policies. Let's break down each layer of the stack in detail — from RAG to production deployment.

Why Do RAG Systems Break and How to Fix It?

RAG (Retrieval-Augmented Generation) looks simple: find relevant documents, put them in context, get an answer. In practice, it fails in several places.

Chunking without overlap. Classic mistake: chunk_size=512, overlap=0. If the answer lies across two chunks, retrieval won't find either with sufficient confidence. Solution: overlap 15–25% of chunk_size, or better yet, sentence-aware splitting with spaCy or NLTK instead of naive character splitting.

Poor embedder. text-embedding-ada-002 is good for general use, but on legal or medical texts, specialized models like E5-large-v2, BGE-M3, or fine-tuned sentence-transformers on domain data outperform it. Recall@5 differences can be 15–25%.

No re-ranking. Vector search optimizes for speed, not relevance. A cross-encoder re-ranker (ms-marco-MiniLM-L-6-v2, bge-reranker-large) after initial retrieval improves top-3 accuracy with acceptable latency (+50–150ms). This is often more impactful than improving the embedding model.

Hybrid search. Dense vectors alone work poorly on exact queries: names, SKUs, codes. BM25 (sparse) finds exact matches but misses semantics. Hybrid via RRF (Reciprocal Rank Fusion) is the optimal compromise. Qdrant, Weaviate, and pgvector 0.7+ support hybrid search natively.

Typical production architecture for a corporate knowledge base
  1. Documents → preprocessing (PyMuPDF, Unstructured)
  2. Chunking → embedding (BGE-M3)
  3. Qdrant (hybrid dense+sparse)
  4. Cross-encoder re-ranking
  5. Context → LLM (vLLM or OpenAI API)
  6. Answer with sources (RAGAS for quality evaluation)

When to Fine-Tune Instead of Prompt Engineering?

Prompt engineering solves ~70% of LLM adaptation tasks for a domain. The remaining 30% require fine-tuning. Three indicators: the model ignores a specific output format even with detailed prompting; the task requires deep knowledge of specialized vocabulary (medicine, law); you need to significantly reduce token costs by replacing a large model with a smaller specialized one.

LoRA and QLoRA are the standard for SFT. LoRA adds trainable low-rank matrices to attention layers. A typical configuration for Llama-3 8B: r=64, lora_alpha=128, target_modules=["q_proj","v_proj","k_proj","o_proj"] yields ~0.8% trainable parameters, training on one A100 40GB. QLoRA adds 4-bit quantization (NF4) and allows fine-tuning 70B models on two A100 40GB, though speed drops by half compared to bf16.

DPO instead of RLHF. Direct Preference Optimization requires only (chosen, rejected) pairs, not scalar reward signals. DPOTrainer from the trl library (Hugging Face) implements it in a few dozen lines.

Common mistake. A dataset of 500 examples, 5 epochs, validation loss 0.8 — seems fine. But on test, the model degrades on general instructions. Cause: catastrophic forgetting. Solution: add 10–20% general instruction-following examples (Alpaca, FLAN) to the training set to preserve original capabilities.

How to Choose a Base Model: 8B or 70B?

Model Parameters Strengths Context
Llama-3.1 8B 8B Quality/speed balance 128k
Llama-3.1 70B 70B Complex reasoning 128k
Mistral 7B / Mixtral 8x7B 7B / 47B Efficiency for size 32k
Qwen2.5 72B 72B Code, multilingual 128k
Gemma 2 27B 27B Open license 8k

For most tasks, fine-tuning an 8B model is sufficient. 70B is needed when deep reasoning is required or the 8B baseline does not reach the required quality even after fine-tuning. Inference cost for Llama-3 8B via vLLM on A100 is efficient; the exact cost depends on volume.

What Does PagedAttention Bring to Production?

vLLM is the first choice for serving open-source models. PagedAttention is the key technical innovation: KV-cache is managed like virtual memory in an OS, without fragmentation. This yields 2–4x higher throughput compared to naive HuggingFace Transformers inference. The vLLM documentation confirms that continuous batching and PagedAttention are the standard for high-load LLM services.

Typical numbers on A100 80GB for Llama-3 8B (bf16): 400–600 req/s, P50 latency 200–400ms, P99 latency 600–900ms at concurrency 64. For 70B on two A100 with tensor parallelism: 80–120 req/s, P99 latency 1.5–2.5s. AWQ or GPTQ quantization reduces memory consumption by 2x with quality loss within 1–3%.

Multi-Agent Systems

Agents are LLMs with access to tools: search, code execution, API calls, database interaction. Common patterns:

  • ReAct (Reason + Act): the model reasons → chooses a tool → observes the result → reasons again. LangChain and LlamaIndex implement it out of the box.
  • Multi-agent orchestration: multiple specialized agents with a coordinator on top. Example: coordinator → researcher (search + summarization) → coder (code generation and execution) → critic (verification). Tools: AutoGen (Microsoft), CrewAI, custom implementation on LangGraph.

In production, agent systems are non-deterministic. Essential: guardrails, step limits, logging of each step, human-in-the-loop for critical actions.

How We Work: Stages, Timeline, Deliverables

Stage Duration What You Get
Audit and data collection 1–2 weeks Eval dataset of 100+ examples, task formalization
Baseline (prompt + RAG) 1–2 weeks Working prototype, quality metrics
Fine-tuning (if needed) 2–4 weeks Trained model, LoRA weights, model card
Deployment and monitoring 1–2 weeks vLLM server, Grafana + Prometheus
Documentation and training 1 week API documentation, team training

What Is Included

We deliver:

  • Technical documentation (model card, configs, deployment instructions)
  • Access to infrastructure (code repository, trained weights)
  • 1 month of post-deployment support (consultations, bug fixes)
  • Customer team training (2–3 sessions on system operation)

Timeline: basic RAG prototype — 1–2 weeks. Fine-tuning with customer data — 3–6 weeks (including data preparation). Production system with monitoring and retraining — 2–4 months. Cost is calculated individually based on data volume, model complexity, and infrastructure requirements.

We guarantee the quality of the final model with performance benchmarks and ongoing monitoring. Our engineers have hands‑on experience with dozens of production LLM systems.

Want to evaluate your project? Leave a request — we will prepare a preliminary summary within 1–2 business days. Or get a consultation on choosing the approach: RAG, fine-tuning, or hybrid — we will tell you what works best for you. Contact us to discuss your LLM development needs. Schedule a free consultation today.