Data engineering

From RSS to JSON: build a feed pipeline you can trust

By Technical Dost · Published

A practical guide to collecting RSS and Atom feeds, choosing stable IDs, handling missing fields and keeping repeat runs useful.

Illustration of feed cards flowing into an organized data table

Start with a feed, not a page scraper

Suppose you want a daily reading queue from a handful of engineering blogs. Before writing selectors for every website, check whether those publishers offer a public RSS or Atom feed. A feed gives you a defined syndication format and a smaller surface to process than a complete page.

RSS 2.0 uses XML and groups items inside a channel. An item may contain a title, link, description, identifier and publication date, but most item fields are optional. Atom has its own structure. Treat them as separate input formats that your parser maps into a common record.

Sources: RSS 2.0 specification · Atom format: RFC 4287

Design the record your workflow actually needs

For a reading queue, begin with source URL, item ID, title, destination URL, publication time and collection time. Keep missing values explicit. A missing date is not the time your collector ran, and a short description is not necessarily the full article.

Keep the original item alongside the normalized fields when debugging matters. That makes it easier to tell whether a blank field came from the publisher or from your transformation. The example below is an application-level shape, not a promised output schema from any particular tool.

Example normalized record with deliberately missing fields
{
  "sourceUrl": "https://example.com/feed.xml",
  "sourceId": "post-42",
  "title": "A practical engineering note",
  "url": "https://example.com/posts/42",
  "publishedAt": null,
  "summary": null
}

Separate collection from deduplication

Fetching the same feed tomorrow may return many of today's items again. Keep persistent state if your destination should receive only new entries. Prefer the publisher's stable item identifier when present; Atom defines a permanent ID, and RSS provides a guid field.

For your own store, pair the source with the item ID. If an identifier is absent, define and document a fallback using the destination URL. Keep an update policy too: the same item can arrive again with a corrected title or description. Deduplication should not silently discard changes you care about.

Choose an output format deliberately

JSON is a useful handoff when the next stage needs nested values. CSV is convenient for inspecting a flat table, but arrays and embedded markup need a defined transformation. Review a small sample before wiring the result into notifications or a report.

If you want to inspect this step without maintaining a feed parser, our RSS Feed Scraper guide explains the public-feed input and dataset output. It collects a feed on each run; new-items-only behaviour still needs an appropriate monitoring or deduplication step.

Sources: Apify dataset formats and exports

Test the quiet failures

Use fixtures with no date, no description, repeated IDs and an empty feed. Record fetch failures separately from a successful fetch with zero items. Decide whether one broken source should stop the whole batch or leave the healthy sources available.

Finally, make a repeat run part of your acceptance check. A useful pipeline produces a result you can explain: which sources succeeded, what changed and why an item was included. Keep attribution and publisher links attached to the material as it moves downstream.

Help us make useful tools easier to find

Allow Google Analytics cookies to measure page visits and clicks to our app stores and Apify. Advertising tracking stays off. Change your choice anytime in the footer. Google privacy policy.