Parser data auto-filling: mapping, moderation, publication

Our company is engaged in the development, support and maintenance of sites of any complexity. From simple one-page sites to large-scale cluster systems built on micro services. Experience of developers is confirmed by certificates from vendors.

Development and maintenance of all types of websites:

Informational websites or web applications
Business card websites, landing pages, corporate websites, online catalogs, quizzes, promo websites, blogs, news resources, informational portals, forums, aggregators
E-commerce websites or web applications
Online stores, B2B portals, marketplaces, online exchanges, cashback websites, exchanges, dropshipping platforms, product parsers
Business process management web applications
CRM systems, ERP systems, corporate portals, production management systems, information parsers
Electronic service websites or web applications
Classified ads platforms, online schools, online cinemas, website builders, portals for electronic services, video hosting platforms, thematic portals

These are just some of the technical types of websites we work with, and each of them can have its own specific features and functionality, as well as be customized to meet the specific needs and goals of the client.

Showing 1 of 1All 2062 services
Parser data auto-filling: mapping, moderation, publication
Medium
~5 days
Frequently Asked Questions

Our competencies:

Development stages

Latest works

  • image_website-b2b-advance_0.webp
    B2B ADVANCE company website development
    1361
  • image_web-applications_feedme_466_0.webp
    Development of a web application for FEEDME
    1253
  • image_websites_belfingroup_462_0.webp
    Website development for BELFINGROUP
    958
  • image_ecommerce_furnoro_435_0.webp
    Development of an online store for the company FURNORO
    1190
  • image_crm_enviok_479_0.webp
    Development of a web application for Enviok
    931
  • image_bitrix-bitrix-24-1c_fixper_448_0.webp
    Website development for FIXPER company
    949

Imagine an online store with 50,000 products, data supplied by three partners in different formats. Direct import within a day leads to 15% duplicates and broken links — the site loses search positions. We developed a system that eliminates such issues through a pipeline: parser → normalization → moderation → publication. But simply copying data from a parser into the database is naive and doesn't scale. Without intermediate processing, you risk junk: duplicates, broken images, invalid prices. Additionally, each supplier sends data in their own format: one uses JSON with nested fields, another uses XML with attributes. We need to unify everything into a single structure, check quality, and only then publish. This auto-filling website method requires systematic approach, otherwise web scraping turns into chaos. Below we break down key nodes and implementation.

What problems arise with direct import?

Direct import from parser to site database is bad practice. Typical errors:

  • Duplicates — identical products with different IDs due to repeated parsing.
  • Broken images — links to external resources that were deleted.
  • Invalid prices — 0 or 999,999,999 rubles from source errors.
  • Different structure — each supplier has their own format: article in article or sku.

We solve this with an intermediate queue with validation and normalization, reducing manual correction by 80%.

Why is an intermediate queue needed?

The queue buffers data and applies business logic: deduplication, format transformation, enrichment from external APIs. Without a queue, a parser failure can dump junk into your CMS. Using a queue reduces publication errors by 5x compared to direct import, processing up to 10,000 records per minute. Our solution, built over 7+ years of experience with 150+ projects, ensures reliability.

How deduplication is implemented?Hashing on title + sku + supplier_id achieves 99.9% accuracy. Duplicates are marked and not published; updates replace old values.

How to configure field mapping for any source?

Each source has its own structure. We use configuration based on JSONPath, allowing field correspondence without code changes.

{
  "source": "supplier_catalog",
  "mappings": {
    "title": "$.name",
    "description": "$.full_description",
    "price": "$.price_rub",
    "category": { "field": "$.category_id", "transform": "category_map" },
    "images": "$.photos[*].url",
    "sku": "$.article"
  },
  "category_map": {
    "1": "electronics",
    "2": "clothing",
    "15": "home-garden"
  }
}

This reduces adaptation time to hours. For CMS import, we integrate via API ensuring seamless field mapping.

How are images processed?

Images are downloaded, optimized, and uploaded to our own storage:

async def process_image(url: str, product_id: int) -> str:
    async with httpx.AsyncClient() as client:
        resp = await client.get(url, timeout=30)
    img = Image.open(BytesIO(resp.content))
    img = img.convert('RGB')
    img.thumbnail((1200, 1200), Image.LANCZOS)
    output = BytesIO()
    img.save(output, 'WEBP', quality=85)
    s3_key = f'products/{product_id}/{uuid4()}.webp'
    s3.put_object(Bucket=BUCKET, Key=s3_key, Body=output.getvalue())
    return f'https://cdn.example.com/{s3_key}'

WebP reduces file size by 30–50% without quality loss, and our CDN caching with 7-day TTL cuts LCP by 40%.

Quality control and publication strategies

Data undergoes three-level validation:

  • Required fields: name, price, at least one photo.
  • Price range: 1 to 1,000,000 units.
  • Description: at least 50 characters.
  • Images: accessible, width ≥ 300 px.

Records failing go to review_required status.

Three publication strategies:

Strategy Publication time Error risk Manual work
Auto-publication Seconds Medium (trusted source) No
Draft Hours/days Low (editor checks) Yes
Diff-update Seconds Low (only changes) No

Choice depends on source reliability and data criticality. This saves up to $20,000 annually for high-volume catalogs.

Process overview

  1. Analysis — study parser structure and CMS fields.
  2. Design — develop mapping and queue schema.
  3. Implementation — write processor in Python, validator with rules, integrate with CMS via API.
  4. Testing — run on 10,000 historical records, check edge cases.
  5. Deployment — deploy to production, configure error monitoring.

What's included

  • Analysis of parser structure and CMS fields
  • Development of mapping schema
  • Configuration of intermediate queue and validation
  • Integration via CMS API
  • Configuration and support documentation
  • Training editors on queue and moderation workflow
  • Technical support during launch

Timing

Complexity Time
One source, basic validation 5–8 days
Multiple sources, UI mapping, moderation 15–20 days

With 7+ years in the field and 150+ successful projects, we ensure reliable parser CMS integration. Contact us for a consultation to evaluate your project — we'll select the optimal architecture for automated content filling.

CMS development: solving real editorial bottlenecks, not installing plugins

A news publisher had a WordPress site with 5 editors. Every article required 15 minutes of manual formatting because the WYSIWYG mangled pasted text. After 6 months, the database had 12 different font sizes and 7 custom colors. The redesign would cost $30k just to clean up the mess — and no one would admit it.

We develop content management systems (CMS) that prevent this from day one. Instead of free-form <textarea> hell, we design structured content models, custom WYSIWYG editors using ProseMirror, and media libraries that offload to S3+CDN within two sprints. This is CMS development without shortcuts.

When is headless CMS justified and when not?

Headless CMS (Strapi, Contentful, Sanity) decouples content management from frontend rendering — the API serves content to any client: website, mobile app, smart display. You get omnichannel delivery and a React/Vue frontend that never touches the admin panel. But if your editors need “save and see” preview and you have no separate frontend team, headless costs extra: you must build a preview layer or use a service like Vercel’s preview deployments.

Sanity customises Studio down to the field level — each field is a React component you can replace. Portable Text (its rich content format) ports to any renderer via custom serializers. For complex editorial workflows with multiple authors, Sanity is the best choice. Contentful offers stable cloud infrastructure with a marketplace of extensions, but monthly bills scale with content volume — typical enterprise plans are $500–$2,000/month. Strapi is self-hosted, open source, with a TypeScript API and custom fields via plugins, but you manage the hosting and backups.

Traditional CMS (WordPress, Craft CMS) works when editors need a familiar admin UI and the frontend is rendered server-side. Craft CMS provides Matrix fields, flexible entry structures, and built-in localization — it’s a professional tool for content teams that need granular permissions and versioning.

How do we build a WYSIWYG editor that doesn’t break layout?

The editor is the most complex component — not a <textarea>. The sweet spot is Tiptap, built on ProseMirror. Every element (headings, lists, tables, code blocks, images) is an extension. Collaborative editing via Yjs works out of the box. Lexical (Meta) is more performant (>60fps typing on mobile) but harder to extend. TinyMCE is a corporate standard at 300KB bundle, but it generates dirty HTML on paste — inline styles, nested <span>, &nbsp; everywhere.

The root cause: pasting from Word. font-family, mso-* properties, empty <span> tags — all leak into the page unless you sanitize. We configure ProseMirror’s pasteRule with DOMPurify to strip everything except allowed tags. Result: clean, semantic HTML that survives a redesign without manual cleanup. Editors save 2–4 hours per week per person.

Media library: from upload to CDN with transformation

Saving files to the server disk is the classic mistake. The disk fills, scaling fails, and CDN becomes impossible. The correct pipeline: upload to S3-compatible storage (AWS S3, Cloudflare R2, MinIO) → CDN (CloudFront, Cloudflare) → on‑the‑fly transformations.

Imgproxy or Thumbor generate any size and format dynamically: https://img.example.com/resize:800:600/format:webp/plain/s3://bucket/photo.jpg. The original lives once, derivatives never occupy disk. Cloudflare Images costs $5 per 100k images, including transformations. Video uploads use Cloudflare Stream or Mux — encode to HLS, adaptive streaming for any bandwidth. Without this, a 1080p video (500MB) loads entirely before play, causing a 5–8 second delay on 3G.

What’s included in media library development

Component Technology Timeline (weeks)
Upload and storage in S3 AWS SDK / MinIO 1–2
Image transformations Imgproxy / Thumbor 1–2
Video streaming Cloudflare Stream / Mux 1–2
Upload and sorting UI React + @dnd-kit/sortable 1–3
Migration of existing files Custom script 0.5–1

Why structured content outperforms free-form HTML

Free-form WYSIWYG leads to chaos in a year: 7 font sizes, 12 colors, random margins. Redesign requires manual cleanup of thousands of posts. Structured content stores “what” instead of “how”: not <p style="font-size:24px; color:red">Important!</p>, but a callout block with variant: warning. The CMS stores the structure; the frontend decides rendering. Sanity Portable Text, Contentful Rich Text, and Strapi Dynamic Zones all follow this pattern — and it reduces rework by 70% during redesigns.

Typical editorial time savings with structured content
  • A news site with 50 articles per week: editors save 10 hours/week on formatting.
  • A corporate portal with 1000 existing pages: migration from free-form to structured content takes 3–5 days, cutting page load by 40% (cleaner HTML).

Work process

  1. Analysis of editorial workflows — who edits, how often, what content (articles, landing pages, product data), whether localization is needed.
  2. CMS selection — based on scenarios, not trends. We compare headless vs traditional with a weighted matrix.
  3. Content model design — record types, fields, relationships, validation rules.
  4. Implementation — frontend integration, editor customization, media library, previews.
  5. Testing — real‑world scenarios: paste from Word, upload 100+ files simultaneously, load test the API (200 req/s target).
  6. Deployment and documentation — editor guide (text + video), API description, access credentials, 1 month support.

Timelines and budget

Type of work Timeline Budget
Integration of headless CMS (Strapi/Sanity) into existing Next.js project 2–5 weeks Discussed individually
Custom WYSIWYG editor with Tiptap and specific blocks 2–4 weeks Discussed individually
Media library with S3 + transformations 1–3 weeks Discussed individually
Full CMS system from scratch 4–10 weeks Discussed individually

Budget is calculated individually after an audit. Client examples: a mid‑sized media site saved $40k/year by eliminating manual formatting; an e‑commerce platform reduced time‑to‑publish by 60% with a headless Sanity setup. Contact us for a free project estimate.

What you get after delivery

  • Working CMS with configured access rights (admin, editor, reviewer)
  • Full content model documentation and API reference
  • Editor training documentation (text + video)
  • Code covered by tests (PHPUnit for Laravel, Jest for JS)
  • 1 month post‑launch support with SLA

Our experience and guarantees

Over 40 completed CMS projects — from small editorial sites to enterprise media portals with 200k daily unique visitors. We use licensed tools (Sentry for error monitoring, SonarCloud for code quality) and guarantee zero critical bugs at launch. All code is version‑controlled and deployable via CI/CD.

For your specific needs, contact us to discuss requirements. We’ll provide a technical proposal within 2 business days.