24HoursNews preview
man 24-hour-news
$ man 24-hour-news

24HoursNews

Work completed at Astavision Infosys

A TypeScript/Express news aggregation backend and Next.js frontend, with MySQL storage, scheduled ingestion, source attribution and searchable categories.

2025 · News aggregation and scheduled ingestion
TypeScriptExpressMySQLNext.jsDrizzle
highlights
$ cat HIGHLIGHTS.md
  • ├─ Publisher-specific ingestion and scheduled jobs
  • ├─ URL-based duplicate prevention in MySQL
  • └─ REST APIs for categories, publishers and search
README.md markdown

Problem

News arrives from publishers with different page structures and date formats. This project brings those sources into a common browsing interface while keeping publisher attribution and links to the original articles.

The repository identifies the backend as work developed at Astavision Infosys.

Approach and architecture

Publisher-specific scrapers feed a TypeScript/Node.js service. That service normalizes article data and writes it to MySQL. Express REST APIs expose the data to a Next.js frontend, including category, publisher and text-search filters.

The current publisher adapters fetch HTML with Axios and parse it with JSDOM. A node-cron scheduler runs ingestion and cleanup in production.

Database decisions

The original implementation used Prisma. The current public source uses Drizzle with MySQL, retaining the existing table and constraint names.

The article URL has a unique constraint. Inserts handle that conflict to avoid saving the same source URL repeatedly. This is URL-level duplicate prevention, rather than an assertion that equivalent stories from different publishers are detected.

Publisher relationships and source URLs travel through the API response so the interface can show where an article came from.

Scheduling and writes

The current scheduler skips a tick if its previous job is still running in the same process. Category ingestion writes are queued within a scraper service instance; cleanup and insertion use a database transaction, with retries for MySQL deadlock errors.

These are practical controls within this application, not coordination across multiple servers.

Engineering lessons

Scraping is only the start of an ingestion pipeline. Date parsing, changing page selectors, duplicate handling and failed writes all affect the data users see. Separating publisher adapters from storage and API formatting makes those responsibilities easier to follow.

Source evidence

The previously listed demo is currently unavailable; the source remains available above.

Let’s talk engineering.

I’m currently open to backend-focused Software Engineer opportunities, including remote roles and relocation. Project enquiries and open-source collaborations are welcome too.

Write to me

Elsewhere