Automated Stock Collection from Supplier Storefronts into 1C-Bitrix
Situation: the supplier provides no API, sends price lists once a week by email, but real stock changes daily. A customer places an order — the item is out of stock. We solve this by parsing the supplier's storefront and updating the CATALOG_QUANTITY field in 1C-Bitrix. Parsing isn't always HTML: often stock data sits in JSON variables on the page (window.__PRODUCT_DATA__). We use CURL and regex — faster than a headless browser and server-friendly. To avoid blocking, we apply random User-Agents and proxy rotation. Over 5 years on the market, we've automated stock collection for catalogs from 500 to 50,000 SKUs. Each project gets individual tuning — from parsing method selection to zero-stock handling logic. Request a parsing development — and we'll analyze your supplier's storefront for free.
How We Parse Stock from Supplier Websites
On the supplier's site, stock can be presented in various ways: numeric value (in stock: 47 pcs) — we directly parse the number; availability status (in stock, pre-order, none) — map to 0/1/999; multiple warehouses — we sum or take the nearest. Sometimes stock is hidden in JS variables — we search with regex in the page body, faster than a headless browser. If the supplier provides an API (rare), we connect via REST.
Why SKU Mapping Is the Tightest Bottleneck
A key stage. Without reliable mapping, parsing is useless. Options:
- Supplier SKU: add a
SUPPLIER_SKU property to the infoblock. While parsing, we search for the element with that value via CIBlockElement::GetList() with a property filter.
- XML_ID: if products were previously imported from the supplier's price list, XML_ID may match their internal ID.
- EAN/barcode: a universal option for branded products.
For large catalogs (10,000+ SKUs), filtering by property via ORM is slow. Better to build a reverse mapping supplier_sku → element_id in Redis or a custom table and update it on catalog changes.
How Stock Updates Work in Bitrix
According to 1C-Bitrix documentation, updating product quantity is done via the CCatalogProduct::Update method.
CCatalogProduct::Update($elementId, [
'QUANTITY' => $parsedQty,
'QUANTITY_RESERVED' => 0,
]);
If the store uses warehouses (b_catalog_store), we update via CCatalogStoreProduct::Update() with the STORE_ID specified.
When updating only the quantity, do not touch ACTIVE — otherwise you'll lose manual activity edits. Perform a separate UPDATE of only the needed field.
Handling Zero Stock Without Losing Sales
Don't automatically hide a product at zero stock — the supplier might restock the next day. The correct gentle logic:
- Quantity = 0 → product remains active, but gets a 'pre-order' flag.
- Quantity = 0 for more than N days → notify manager, manual decision.
- Product not found on supplier site 3+ times in a row → flag
SUPPLIER_DISCONTINUED.
We implement flags via infoblock properties or Highload-block fields.
Step-by-Step Process
-
Supplier storefront analysis. Determine parsing method (HTML, JSON variables, API), prepare selectors.
-
Parser development. PHP script with CURL, regex, optionally Guzzle. For blocking protection, use random User-Agents and proxies.
-
SKU mapping setup. Link supplier SKUs to Bitrix product IDs. For catalogs 10,000+ SKUs, use Redis.
-
Bitrix integration. Agent or cron task, update stock via
CCatalogProduct::Update(). Test on a sample set of products.
-
Zero stock handling logic. Automatic flags, manager notifications. Test scenarios: product disappears, appears, price changes.
-
Monitoring and support. After launch, parser runs stably; if supplier site structure changes, adaptation takes no more than a couple of hours.
What's Included in the Work
| Component |
Description |
| Supplier storefront analysis |
Determine parsing method (HTML, JSON variables, API), prepare selectors |
| Parser development |
PHP script with CURL, regex, optionally Guzzle |
| SKU mapping |
Link supplier SKUs to Bitrix product IDs |
| Bitrix integration |
Agent or cron task, update stock via CCatalogProduct::Update() |
| Zero stock handling logic |
Automatic flags, manager notifications |
| Documentation and support |
Mapping scheme, instructions for adding a supplier, 2 weeks free support |
Timeline
| Stage |
Duration |
| Supplier site analysis, parsing method selection |
2–4 hours |
| Parser development |
1–2 days |
| SKU mapping setup |
4–8 hours |
| Bitrix update logic + zero stock handling |
4–8 hours |
| Schedule and monitoring setup |
2–4 hours |
Total: 3–5 working days for one supplier. Each additional supplier — +1–2 days (different site structures).
We guarantee that after launch the parser will run stably, and if the supplier's site structure changes, adaptation will take no more than a couple of hours. Time savings on manual stock processing: up to 80% per month. Get a free assessment of your project — just contact us. Get a consultation right now.
Parser Development for 1C-Bitrix: Where to Start?
XMLReader, not SimpleXML — the choice of tool determines the project's fate. SimpleXML loads the entire XML into memory, and with an 800 MB supplier file, PHP will crash with a fatal error on a 512 MB limit. XMLReader processes streamingly, node by node, consuming 20–30 MB — 30 times more efficient. This detail starts any parser development for Bitrix. With over 10 years of Bitrix development and 50+ parser projects delivered, we know the pitfalls. Contact us to start your parser development today.
What Problems Does Parsing Solve?
- Primary catalog filling — 15,000 cards with descriptions, characteristics, photos. Manually, that's three months of content manager work; a parser takes a week with debugging.
- Competitor price monitoring — collecting data from Ozon, Wildberries, competitor sites. A competitor drops the price on a hot item — you find out in two hours, not two weeks.
- Supplier aggregation — five price lists in different formats (CSV with CP1251, XML in CommerceML, Excel with merged cells) become a single catalog with a unified property system.
- Card enrichment — pulling characteristics, instructions, 3D models from manufacturer sites. Without this, a product card is an SEO empty shell.
- Assortment update — products missing from the supplier feed are deactivated via
CIBlockElement::Update($ID, ['ACTIVE' => 'N']). New ones are created. The catalog stays synchronized.
What Tools Do We Use in Parser Development?
Static websites — PHP (Goutte, Symfony DomCrawler) or Python (Scrapy, lxml). Speed: 50–100 pages/sec. Sufficient for catalogs without JS rendering.
SPA and dynamic websites — Puppeteer or Playwright. Infinite scroll, AJAX filters, lazy-load images — headless browser handles it all. Speed drops to 1–10 pages/sec, but there is no alternative: data exists only after JavaScript execution.
Supplier files:
- Excel (XLS, XLSX) — PhpSpreadsheet. Beware of merged cells and formulas — they break automatic mapping.
- CSV —
fgetcsv() with correct encoding. Suppliers love CP1251, BOM in UTF-8, and semicolons instead of commas. All need detection and handling.
- XML/YML — XMLReader for large files, SimpleXML for feeds up to 50 MB.
- CommerceML — standard exchange format with 1C. We parse
import.xml and offers.xml, map to information block structure.
API — Supplier REST endpoints, marketplace APIs (Ozon Seller API, Wildberries API). We work within rate limits, handle pagination.
How Is the Auto-Population Pipeline Structured?
Four stages. Each can break in its own way.
-
Collection. Parser crawls sources via cron schedule. Raw data goes to an intermediate table — not directly into b_iblock_element. Log everything: pages visited, elements parsed, where we got 403 or timeout. Without logs, debugging a parser is like fortune-telling.
-
Normalization. Main work here:
- Clean HTML tags, extra spaces, Unicode garbage
- Units: "mm" → "mm", "millimeters" → "mm", "миллиметр" → "mm"
- Map supplier categories to Bitrix information block sections. One supplier has "Notebooks", another "Notebooks and tablets", third "Laptops" — all into one section
- Deduplication by SKU, EAN/GTIN. One product from three suppliers should not appear three times
-
Load into Bitrix. Via CIBlockElement::Add() for new elements, CIBlockElement::Update() for existing. Images: download, resize via CFile::ResizeImageGet(), convert to WebP. Properties via CIBlockElement::SetPropertyValuesEx(). SEO meta via \Bitrix\Iblock\InheritedProperty\ElementValues. SEF URLs generated from name transliteration.
-
Update. Key point — not overwrite manual edits by content manager. Update only price, stock, activity. Description and photos manually edited are flagged with UF_MANUAL_EDIT property and skipped during import. Products missing from feed are deactivated, not deleted.
Why Is Competitor Price Monitoring Necessary?
A separate subsystem with its own specifics:
| Parameter |
How It Works |
| Frequency |
From once a day to every 2 hours — depends on market volatility |
| Matching |
By SKU, EAN, fuzzy name comparison via Levenshtein distance |
| Storage |
Separate vendor_price_monitor table with history, not information blocks |
| Alerts |
Telegram/email when competitor price deviation exceeds X% |
| Auto-rules |
"Keep price 3% below competitor minimum, but not below cost + 15%" |
Result — dashboard: your product vs competitors, price history, trends. The manager sees where to raise price without losing position, and where to react.
CSV/XML Import Module: Customization for Your Format
For supplier files — custom module with admin panel:
- Configurable mapping: "column B in file → BRAND property of information block"
- Auto-detect encoding (CP1251, UTF-8, UTF-16) via
mb_detect_encoding() with validation
- Download images from URL with queue — to avoid channel saturation
- Incremental update by row hash: row changed — update, no — skip
- Cron schedule, report: created 145, updated 892, errors 3 (with details)
Large files: CSV processed in batches of 1000 rows via fgetcsv() (10 times faster than row-by-row), XML streamed via XMLReader, background execution via Bitrix agent queue — no PHP timeouts.
Legal Aspects to Consider
-
robots.txt — respect it. Crawl-delay — comply.
- Request frequency — 1–2 per second, no more. Don't DDoS someone else's site.
- Manufacturer content — use it. Unique author texts — don't copy.
- Personal data — don't collect.
What Is Included in a Turnkey Parser Development?
| Component |
Description |
| Prototype |
Parser for 1–2 sources in 2–3 days to assess data quality |
| Main parser |
Full data collection from one source (static/dynamic) |
| Bitrix import module |
Normalization, loading, update, mapping admin panel |
| Price monitoring |
If needed — collection and alert system (up to 10 competitors) |
| Documentation |
Architecture description, selector update instructions |
| Support |
3-month guarantee for uninterrupted operation, fix for donor layout changes |
How We Work and Deadlines
- Prototype — parser for 1–2 sources in 2–3 days. Assess data quality, pitfalls (Cloudflare protection, captcha, dynamic loading).
- Development — full pipeline: parser → normalization → import into Bitrix → admin panel for management.
- Testing — run on full catalog volume, check edge cases (empty fields, malformed HTML, broken images).
- Launch — configure cron, error monitoring via Telegram bot.
- Support — competitor changed layout? Update CSS selectors in parser.
| Task |
Deadlines |
| Single site parser (static HTML) |
3–5 days |
| SPA site parser (Puppeteer/Playwright, bypass protection) |
1–2 weeks |
| CSV/XML import module for Bitrix |
1–2 weeks |
| Price monitoring system (5–10 competitors) |
2–4 weeks |
| Comprehensive auto-population system |
4–8 weeks |
| Parser support and adaptation |
by subscription |
Get in touch for a free consultation — we will analyze your data sources and propose the optimal parser architecture. Request a project assessment today and get a fixed deadline. We guarantee stable parser operation and full support throughout the usage period.