Parsing product characteristics for Bitrix catalog population

Our company is engaged in the development, support and maintenance of Bitrix and Bitrix24 solutions of any complexity. From simple one-page sites to complex online stores, CRM systems with 1C and telephony integration. The experience of developers is confirmed by certificates from the vendor.
Showing 1 of 1All 1626 services
Parsing product characteristics for Bitrix catalog population
Medium
~1-2 weeks
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1368
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    956
  • image_bitrix-bitrix-24-1c_development_of_an_online_appointment_booking_widget_for_a_medical_center_594_0.webp
    Development based on Bitrix, Bitrix24, 1C for the company Development of an Online Appointment Booking Widget for a Medical Center
    699
  • image_bitrix-bitrix-24-1c_mirsanbel_458_0.webp
    Development based on 1C Enterprise for MIRSANBEL
    843
  • image_crm_dolbimby_434_0.webp
    Website development on CRM Bitrix24 for DOLBIMBY
    737
  • image_crm_technotorgcomplex_453_0.webp
    Development based on Bitrix24 for the company TECHNOTORGKOMPLEKS
    1086

When populating a 1C-Bitrix catalog, we often encounter: product descriptions exist, but characteristics are empty. Without filled properties (b_iblock_element_property), the smart filter returns no results, faceted search gives empty output. Buyers cannot filter products by required parameters — conversion drops, bounce rates rise. Parsing characteristics is technically more complex than parsing descriptions: you need not just extract text, but recognize the structure "parameter name — value" and correctly place data in infoblock properties. Our comprehensive bitrix characteristics parsing and 1c bitrix product specs parsing service handles HTML specification parsing, normalizes property names, and ensures accurate bitrix product characteristics population — all with a focus on smart filter performance. This process is 10 times faster than manual entry. Typical project cost: from $500 for 1000 SKUs to $2000 for complex cases, yielding up to 80% savings over manual data entry. We can assess your project in one day — just reach out to us. Our company has 5+ years of experience and over 50 successful Bitrix projects, ensuring reliable parsing.

Features of parsing characteristics in different formats

The source specification table can be of several types, each requiring a separate extraction approach:

  • HTML table (<table>) — classic, parsed via XPath (see XPath) //table//tr. The first cell of the row is the name, the second is the value.
  • dl/dt/dd list — often used in modern stores. We parse dt+dd pairs.
  • JSON-LD or schema.org microdata — ideal. Data is already structured, no HTML parsing needed:
    preg_match('/<script type="application/ld+json">(.*?)<\/script>/s', $html, $m);
    $data = json_decode($m[1], true);
    
  • JS variables — data in window.productData or __REDUX_STATE__. Extract via regex.

Normalizing characteristic names

Different sources call the same thing differently: "Weight", "Net weight", "Weight (kg)". Direct mapping to an infoblock property without normalization creates chaos. Our solution — an alias table property_aliases:

CREATE TABLE parser_property_aliases (
    alias VARCHAR(255),
    canonical_name VARCHAR(255),
    property_code VARCHAR(100)
);

During parsing, each found name is looked up in the alias table. If not found — it is logged as "unknown property" for manual review and addition to the dictionary. This ensures data cleanliness even when working with thousands of SKUs. The alias dictionary is a key element that prevents duplication and filter errors.

Property types for the smart filter

Infoblock properties (b_iblock_property) have types: S (string), N (number), L (list), E (element binding). For characteristics we use:

Type Description Example Speed in filter
S Text value Color: red Medium
N Numeric with unit Weight: 1.5 High
L Fixed list of values Brand: Samsung High (indexed)

For the faceted filter, values of type L work faster — they are indexed in b_iblock_element_prop_enum. Creating a value on import:

$propEnum = CIBlockPropertyEnum::GetList([], [
    'PROPERTY_ID' => $propId,
    'VALUE' => $parsedValue
])->Fetch();
if (!$propEnum) {
    CIBlockPropertyEnum::Add(['PROPERTY_ID' => $propId, 'VALUE' => $parsedValue]);
}

Handling units of measurement

Sources provide "10 kg", "10kg", "10 kilograms". A unit parser is needed: split number and unit, normalize unit to standard. Simple regex: /^([\d.,]+)\s*(.*)$/. Numeric property values in Bitrix are stored as strings in b_iblock_element_property.VALUE — units should be moved to a separate property or added to the property CODE (WEIGHT_KG).

Case from our practice: electronics, 15,000 SKUs, 120+ characteristic types

Task: fill properties for the smart filter for laptops, phones, TVs — three infoblocks with different property sets.

Implementation:

  • Parsing from the manufacturer's site via JSON-LD (70% of products) + HTML table (30%)
  • Alias dictionary of 380 entries, collected in the first 3 days of development
  • All numeric characteristics — type N, list ones (brand, color, country) — type L
  • Parallel launch of 5 workers via PHP-CLI, each processing its own category

Result: the smart filter worked correctly for 48 parameters after 2 iterations of debugging the alias dictionary. Conversion from the filter increased by 15%. Economy: from $500 to $2000 per 1000 SKUs, depending on complexity.

Process: step-by-step plan

  1. Analysis of source characteristic structure — determine specification format and data volume. 4–8 hours.
  2. Parser development — write extraction module for your source type. 2–3 days.
  3. Creation of alias dictionary — collect and normalize characteristic names. 1–2 days.
  4. Configuration of infoblock properties — create or adjust properties for the filter. 1 day.
  5. Data import and debugging — load characteristics, check types. 1–2 days.
  6. Verification of smart filter operation — test filtering by all parameters. 4–8 hours.

What’s included

  • Full documentation of the alias dictionary and parser configuration
  • Access to the parser script with source code (can be integrated into your Bitrix environment)
  • Training for managers on how to maintain the alias dictionary and handle new sources
  • Post-launch support for 1 month, including fixes and adjustments

What is included in the work

Stage Description Time
Source characteristic structure analysis Determine specification format and data volume 4–8 hours
Parser development Write extraction module for your source type 2–3 days
Alias dictionary creation Collect and normalize characteristic names 1–2 days
Infoblock property configuration Create or adjust properties for the filter 1 day
Data import and debugging Load characteristics, check types 1–2 days
Smart filter operation check Test filtering by all parameters 4–8 hours

We provide documentation, access to the parser, training for managers, and post-launch support. Total: 7–12 business days. This is one of the most labor-intensive catalog population tasks, but we perform it turnkey with a guaranteed result. Certified specialists with 5+ years of experience and over 50 successful Bitrix catalog population projects.

SQL schema of alias table (example)
CREATE TABLE parser_property_aliases (
    id INT AUTO_INCREMENT PRIMARY KEY,
    alias VARCHAR(255) NOT NULL,
    canonical_name VARCHAR(255) NOT NULL,
    property_code VARCHAR(100) NOT NULL,
    created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP
);

Contact us for an assessment of your project — we will analyze your source structure and propose an optimal solution. Get a consultation right now and find out how characteristics parsing can reduce your online store's operational costs.

Parser Development for 1C-Bitrix: Where to Start?

XMLReader, not SimpleXML — the choice of tool determines the project's fate. SimpleXML loads the entire XML into memory, and with an 800 MB supplier file, PHP will crash with a fatal error on a 512 MB limit. XMLReader processes streamingly, node by node, consuming 20–30 MB — 30 times more efficient. This detail starts any parser development for Bitrix. With over 10 years of Bitrix development and 50+ parser projects delivered, we know the pitfalls. Contact us to start your parser development today.

What Problems Does Parsing Solve?

  • Primary catalog filling — 15,000 cards with descriptions, characteristics, photos. Manually, that's three months of content manager work; a parser takes a week with debugging.
  • Competitor price monitoring — collecting data from Ozon, Wildberries, competitor sites. A competitor drops the price on a hot item — you find out in two hours, not two weeks.
  • Supplier aggregation — five price lists in different formats (CSV with CP1251, XML in CommerceML, Excel with merged cells) become a single catalog with a unified property system.
  • Card enrichment — pulling characteristics, instructions, 3D models from manufacturer sites. Without this, a product card is an SEO empty shell.
  • Assortment update — products missing from the supplier feed are deactivated via CIBlockElement::Update($ID, ['ACTIVE' => 'N']). New ones are created. The catalog stays synchronized.

What Tools Do We Use in Parser Development?

Static websites — PHP (Goutte, Symfony DomCrawler) or Python (Scrapy, lxml). Speed: 50–100 pages/sec. Sufficient for catalogs without JS rendering.

SPA and dynamic websites — Puppeteer or Playwright. Infinite scroll, AJAX filters, lazy-load images — headless browser handles it all. Speed drops to 1–10 pages/sec, but there is no alternative: data exists only after JavaScript execution.

Supplier files:

  • Excel (XLS, XLSX) — PhpSpreadsheet. Beware of merged cells and formulas — they break automatic mapping.
  • CSV — fgetcsv() with correct encoding. Suppliers love CP1251, BOM in UTF-8, and semicolons instead of commas. All need detection and handling.
  • XML/YML — XMLReader for large files, SimpleXML for feeds up to 50 MB.
  • CommerceML — standard exchange format with 1C. We parse import.xml and offers.xml, map to information block structure.

API — Supplier REST endpoints, marketplace APIs (Ozon Seller API, Wildberries API). We work within rate limits, handle pagination.

How Is the Auto-Population Pipeline Structured?

Four stages. Each can break in its own way.

  1. Collection. Parser crawls sources via cron schedule. Raw data goes to an intermediate table — not directly into b_iblock_element. Log everything: pages visited, elements parsed, where we got 403 or timeout. Without logs, debugging a parser is like fortune-telling.

  2. Normalization. Main work here:

    • Clean HTML tags, extra spaces, Unicode garbage
    • Units: "mm" → "mm", "millimeters" → "mm", "миллиметр" → "mm"
    • Map supplier categories to Bitrix information block sections. One supplier has "Notebooks", another "Notebooks and tablets", third "Laptops" — all into one section
    • Deduplication by SKU, EAN/GTIN. One product from three suppliers should not appear three times
  3. Load into Bitrix. Via CIBlockElement::Add() for new elements, CIBlockElement::Update() for existing. Images: download, resize via CFile::ResizeImageGet(), convert to WebP. Properties via CIBlockElement::SetPropertyValuesEx(). SEO meta via \Bitrix\Iblock\InheritedProperty\ElementValues. SEF URLs generated from name transliteration.

  4. Update. Key point — not overwrite manual edits by content manager. Update only price, stock, activity. Description and photos manually edited are flagged with UF_MANUAL_EDIT property and skipped during import. Products missing from feed are deactivated, not deleted.

Why Is Competitor Price Monitoring Necessary?

A separate subsystem with its own specifics:

Parameter How It Works
Frequency From once a day to every 2 hours — depends on market volatility
Matching By SKU, EAN, fuzzy name comparison via Levenshtein distance
Storage Separate vendor_price_monitor table with history, not information blocks
Alerts Telegram/email when competitor price deviation exceeds X%
Auto-rules "Keep price 3% below competitor minimum, but not below cost + 15%"

Result — dashboard: your product vs competitors, price history, trends. The manager sees where to raise price without losing position, and where to react.

CSV/XML Import Module: Customization for Your Format

For supplier files — custom module with admin panel:

  • Configurable mapping: "column B in file → BRAND property of information block"
  • Auto-detect encoding (CP1251, UTF-8, UTF-16) via mb_detect_encoding() with validation
  • Download images from URL with queue — to avoid channel saturation
  • Incremental update by row hash: row changed — update, no — skip
  • Cron schedule, report: created 145, updated 892, errors 3 (with details)

Large files: CSV processed in batches of 1000 rows via fgetcsv() (10 times faster than row-by-row), XML streamed via XMLReader, background execution via Bitrix agent queue — no PHP timeouts.

Legal Aspects to Consider

  • robots.txt — respect it. Crawl-delay — comply.
  • Request frequency — 1–2 per second, no more. Don't DDoS someone else's site.
  • Manufacturer content — use it. Unique author texts — don't copy.
  • Personal data — don't collect.

What Is Included in a Turnkey Parser Development?

Component Description
Prototype Parser for 1–2 sources in 2–3 days to assess data quality
Main parser Full data collection from one source (static/dynamic)
Bitrix import module Normalization, loading, update, mapping admin panel
Price monitoring If needed — collection and alert system (up to 10 competitors)
Documentation Architecture description, selector update instructions
Support 3-month guarantee for uninterrupted operation, fix for donor layout changes

How We Work and Deadlines

  1. Prototype — parser for 1–2 sources in 2–3 days. Assess data quality, pitfalls (Cloudflare protection, captcha, dynamic loading).
  2. Development — full pipeline: parser → normalization → import into Bitrix → admin panel for management.
  3. Testing — run on full catalog volume, check edge cases (empty fields, malformed HTML, broken images).
  4. Launch — configure cron, error monitoring via Telegram bot.
  5. Support — competitor changed layout? Update CSS selectors in parser.
Task Deadlines
Single site parser (static HTML) 3–5 days
SPA site parser (Puppeteer/Playwright, bypass protection) 1–2 weeks
CSV/XML import module for Bitrix 1–2 weeks
Price monitoring system (5–10 competitors) 2–4 weeks
Comprehensive auto-population system 4–8 weeks
Parser support and adaptation by subscription

Get in touch for a free consultation — we will analyze your data sources and propose the optimal parser architecture. Request a project assessment today and get a fixed deadline. We guarantee stable parser operation and full support throughout the usage period.