Collecting content from dozens of sources manually takes hours of copying, duplicate checking, and sorting. An automated content aggregator solves this: it parses RSS and APIs, filters duplicates, categorizes, and presents a personalized feed. We design and implement such a pipeline for your project — from architecture to deployment. Over the years we have launched more than 10 aggregators for media and corporate portals, reducing manual content-gathering by 80%. Budget savings reach 60% thanks to automation.
A typical project includes 30–50 RSS feeds, 5–10 API sources, and several websites for scraping. Each source requires a separate adapter with rate-limit and error handling. We use Bull job queues on Redis, allowing parallel processing of up to 100 sources without data loss. On failure, a task automatically retries with exponential backoff up to 3 times. Monitoring via Grafana tracks the success rate of each source.
Pipeline Schema
| Stage |
Tool |
Description |
| Scheduler |
cron |
Triggers collection on a schedule every N minutes |
| Fetcher |
Bull/BullMQ |
Job queue per source with retries |
| Parser |
rss-parser, Cheerio, Playwright |
Extract data from RSS, HTML, SPA |
| Normalizer |
custom code |
Map fields to a unified format |
| Deduplicator |
SimHash, MinHash |
Detect exact and near duplicates |
| Storage |
PostgreSQL |
Primary storage |
| Indexer |
Elasticsearch / Meilisearch |
Full-text search and filtering |
Why Deduplication Is Key
Duplicates appear when the same story is published by multiple sources. We apply three methods:
| Method |
Principle |
Effectiveness |
| Exact URL match |
Check URL uniqueness |
Only identical URLs |
| Title hash |
Hash of normalized title |
Identical titles |
| SimHash / MinHash |
Approximate near-duplicate detection |
Similar texts (configurable threshold) |
SimHash is 3× more effective than exact hashing for detecting similar texts, reducing false positives to 5%. With proper tuning, deduplication filters out 95% of duplicates. For one media project with 50 RSS and 10 API sources, we set up a pipeline that processes 500 articles daily, slashing manual work from 4 hours to 15 minutes. More on SimHash.
from simhash import Simhash
def is_duplicate(text1: str, text2: str, threshold: int = 5) -> bool:
h1, h2 = Simhash(text1.split()), Simhash(text2.split())
return h1.distance(h2) < threshold
What Parsing Tools Do We Use?
RSS parsing is straightforward. The challenge is extracting clean text when scraping:
-
Readability (Mozilla) – strips navigation and ads;
-
Trafilatura (Python) – extracts text with language detection;
-
Playwright – for SPAs requiring full JavaScript rendering.
Each adapter takes 1–2 days to set up, including error handling and rate limits. Parsing speed reaches 10 articles per second per adapter. Readability docs: Readability.
Categorization and Tagging
Automatic classification by topic:
-
Keyword matching: rules like "rouble" → "Finance";
-
ML classification: fastText or BERT-based for multi-label tagging. ML accuracy reaches 90% on labeled data. Processing 1000 articles per minute is a realistic speed for fastText.
Language detection: langdetect (Python) or franc (Node.js).
How Feed Personalization Works
The user manages filters:
- enabled/disabled sources and categories;
- keyword subscriptions;
- negative keywords to exclude topics.
Optionally, algorithmic ranking: collaborative filtering based on reading history of similar users. This boosts engagement by 30% according to our data.
Copyright Compliance
The aggregator shows only previews (lead + link), respects robots.txt and rate limits. The fair use model allows snippets but not full republication. We always credit the source and author.
What Is Included in the Work
- Source analysis and architecture design;
- Pipeline implementation: parsing → normalization → deduplication → storage → delivery;
- Indexing setup in Elasticsearch/Meilisearch;
- Categorization (rules or ML);
- Personalization (filters and ranking);
- API and admin panel documentation;
- Testing and load testing;
- Deployment with monitoring (Grafana, Sentry).
Timeline
| Stage |
Duration |
| MVP (10–20 RSS, feed, search, categories) |
4–6 weeks |
| Full feature set (ML, scraping, personalization, API) |
3–5 months |
Cost is determined individually after an audit. Order an audit of your sources — we’ll evaluate in one day. We guarantee deadlines and confidentiality. Our team has launched over 10 aggregators. Contact us to discuss your project.
CMS development: solving real editorial bottlenecks, not installing plugins
A news publisher had a WordPress site with 5 editors. Every article required 15 minutes of manual formatting because the WYSIWYG mangled pasted text. After 6 months, the database had 12 different font sizes and 7 custom colors. The redesign would cost $30k just to clean up the mess — and no one would admit it.
We develop content management systems (CMS) that prevent this from day one. Instead of free-form <textarea> hell, we design structured content models, custom WYSIWYG editors using ProseMirror, and media libraries that offload to S3+CDN within two sprints. This is CMS development without shortcuts.
When is headless CMS justified and when not?
Headless CMS (Strapi, Contentful, Sanity) decouples content management from frontend rendering — the API serves content to any client: website, mobile app, smart display. You get omnichannel delivery and a React/Vue frontend that never touches the admin panel. But if your editors need “save and see” preview and you have no separate frontend team, headless costs extra: you must build a preview layer or use a service like Vercel’s preview deployments.
Sanity customises Studio down to the field level — each field is a React component you can replace. Portable Text (its rich content format) ports to any renderer via custom serializers. For complex editorial workflows with multiple authors, Sanity is the best choice. Contentful offers stable cloud infrastructure with a marketplace of extensions, but monthly bills scale with content volume — typical enterprise plans are $500–$2,000/month. Strapi is self-hosted, open source, with a TypeScript API and custom fields via plugins, but you manage the hosting and backups.
Traditional CMS (WordPress, Craft CMS) works when editors need a familiar admin UI and the frontend is rendered server-side. Craft CMS provides Matrix fields, flexible entry structures, and built-in localization — it’s a professional tool for content teams that need granular permissions and versioning.
How do we build a WYSIWYG editor that doesn’t break layout?
The editor is the most complex component — not a <textarea>. The sweet spot is Tiptap, built on ProseMirror. Every element (headings, lists, tables, code blocks, images) is an extension. Collaborative editing via Yjs works out of the box. Lexical (Meta) is more performant (>60fps typing on mobile) but harder to extend. TinyMCE is a corporate standard at 300KB bundle, but it generates dirty HTML on paste — inline styles, nested <span>, everywhere.
The root cause: pasting from Word. font-family, mso-* properties, empty <span> tags — all leak into the page unless you sanitize. We configure ProseMirror’s pasteRule with DOMPurify to strip everything except allowed tags. Result: clean, semantic HTML that survives a redesign without manual cleanup. Editors save 2–4 hours per week per person.
Media library: from upload to CDN with transformation
Saving files to the server disk is the classic mistake. The disk fills, scaling fails, and CDN becomes impossible. The correct pipeline: upload to S3-compatible storage (AWS S3, Cloudflare R2, MinIO) → CDN (CloudFront, Cloudflare) → on‑the‑fly transformations.
Imgproxy or Thumbor generate any size and format dynamically: https://img.example.com/resize:800:600/format:webp/plain/s3://bucket/photo.jpg. The original lives once, derivatives never occupy disk. Cloudflare Images costs $5 per 100k images, including transformations. Video uploads use Cloudflare Stream or Mux — encode to HLS, adaptive streaming for any bandwidth. Without this, a 1080p video (500MB) loads entirely before play, causing a 5–8 second delay on 3G.
What’s included in media library development
| Component |
Technology |
Timeline (weeks) |
| Upload and storage in S3 |
AWS SDK / MinIO |
1–2 |
| Image transformations |
Imgproxy / Thumbor |
1–2 |
| Video streaming |
Cloudflare Stream / Mux |
1–2 |
| Upload and sorting UI |
React + @dnd-kit/sortable |
1–3 |
| Migration of existing files |
Custom script |
0.5–1 |
Why structured content outperforms free-form HTML
Free-form WYSIWYG leads to chaos in a year: 7 font sizes, 12 colors, random margins. Redesign requires manual cleanup of thousands of posts. Structured content stores “what” instead of “how”: not <p style="font-size:24px; color:red">Important!</p>, but a callout block with variant: warning. The CMS stores the structure; the frontend decides rendering. Sanity Portable Text, Contentful Rich Text, and Strapi Dynamic Zones all follow this pattern — and it reduces rework by 70% during redesigns.
Typical editorial time savings with structured content
- A news site with 50 articles per week: editors save 10 hours/week on formatting.
- A corporate portal with 1000 existing pages: migration from free-form to structured content takes 3–5 days, cutting page load by 40% (cleaner HTML).
Work process
-
Analysis of editorial workflows — who edits, how often, what content (articles, landing pages, product data), whether localization is needed.
-
CMS selection — based on scenarios, not trends. We compare headless vs traditional with a weighted matrix.
-
Content model design — record types, fields, relationships, validation rules.
-
Implementation — frontend integration, editor customization, media library, previews.
-
Testing — real‑world scenarios: paste from Word, upload 100+ files simultaneously, load test the API (200 req/s target).
-
Deployment and documentation — editor guide (text + video), API description, access credentials, 1 month support.
Timelines and budget
| Type of work |
Timeline |
Budget |
| Integration of headless CMS (Strapi/Sanity) into existing Next.js project |
2–5 weeks |
Discussed individually |
| Custom WYSIWYG editor with Tiptap and specific blocks |
2–4 weeks |
Discussed individually |
| Media library with S3 + transformations |
1–3 weeks |
Discussed individually |
| Full CMS system from scratch |
4–10 weeks |
Discussed individually |
Budget is calculated individually after an audit. Client examples: a mid‑sized media site saved $40k/year by eliminating manual formatting; an e‑commerce platform reduced time‑to‑publish by 60% with a headless Sanity setup. Contact us for a free project estimate.
What you get after delivery
- Working CMS with configured access rights (admin, editor, reviewer)
- Full content model documentation and API reference
- Editor training documentation (text + video)
- Code covered by tests (PHPUnit for Laravel, Jest for JS)
- 1 month post‑launch support with SLA
Our experience and guarantees
Over 40 completed CMS projects — from small editorial sites to enterprise media portals with 200k daily unique visitors. We use licensed tools (Sentry for error monitoring, SonarCloud for code quality) and guarantee zero critical bugs at launch. All code is version‑controlled and deployable via CI/CD.
For your specific needs, contact us to discuss requirements. We’ll provide a technical proposal within 2 business days.