Nordic Startup News logo
Nordic Startup News European Startup Intelligence
Research overview

How the briefing pipeline actually works

This page explains the public research flow behind the live startup briefing: source intake, clustering, AI-assisted editorial selection, grounded publication and JDBIN-backed delivery.

01RSS and article feeds are normalized into a shared candidate pool.
02Clusters are shortlisted with AI, but publication still goes through deterministic grounding and storage rules.
03The public site reads already published output, not fresh AI generation on every page view.
What this page is for

Separate the public research story from the raw docs tree

The documentation library is useful when you need full implementation detail. The hero CTA on the live homepage needs a cleaner entry point: what the research system does, why it exists, and where each layer fits in the production path.

Public explanation

Explain the pipeline in product terms without sending a first-click visitor into raw markdown or low-level engineering pages.

Operational clarity

Show where source collection ends, where AI ranking begins, and where the grounded publish gate decides what can go live.

Auditability

Keep a direct bridge from this overview into the deeper docs for storage, ingest, benchmarks, object layout and safety boundaries.

Research model

The live page is a published artifact, not an on-demand AI answer

The homepage should show already published briefing output. AI runs in scheduled ingest or manual trigger flows. The read path serves a published view from the JDBIN-backed chain and its public snapshot, so page traffic does not regenerate the briefing.

Fetch and normalizeWorker pulls feeds, parses article metadata, extracts titles, links, timestamps, summaries, categories and source IDs.
Cluster and dedupeRelated articles collapse into story clusters so the public product does not show one topic as repeated raw headlines.
Shortlist and groundWorkers AI validates event frames, OpenAI selects candidates and drafts editorial copy, and the Worker validates source grounding before publish.
Publish to JDBINAccepted rows are appended as immutable JDBIN/JDBON state in R2. Public reads then come from the published chain and canonical public snapshot.
Live audit path

The production path now exposes both publication state and trace metadata

The current implementation does not stop at “AI wrote something.” Each live run now carries enough metadata to verify whether the public snapshot, research snapshot and selected shortlist all refer to the same run, and whether the model output had to be patched before publication.

`/api/trigger` now returns pipeline audit

The trigger response includes a compact operational summary: raw article count, clustered story count, selected candidate count, OpenAI output count, reconciled final count, published signal count and any backfilled ranks.

Main fieldsselectedCandidateCount, openAiSignalCount, reconciledSignalCount, publishedSignalCount
Read-path contextactiveRunRowCount, dedupedRowCount, filteredOutCount
Gap handlingbackfilledRanks
`/api/data` now returns public audit

The public read path exposes how many rows were available in the active run, how many survived dedupe, how many were editorially ready and what finally became published output in the briefing.

Main fieldssourceRowCount, activeRunRowCount, editorialReadyRowCount
Publish outcomepublishedSignalCount, briefingSignalCount
Trace keyspublicSnapshot.key, objectKey, manifestKey
Entry points

Go deeper from the right layer

AI pipeline breakdown

Detailed implementation view of research intake, shortlist rules, grounded copy generation and publication.

Open pipeline breakdown

JDBIN system

Storage model, R2 object chain, manifest/pointer control, query path and immutable publish architecture.

Open storage model

Ingest system

How scheduled runs, manual trigger flows and Worker ingest write new publishable rows into the active chain.

Open ingest flow

Publishing model

How the product decides what can become public output and why read traffic is separated from write-time AI generation.

Open publishing model

Next step

Use this as the research-facing landing page

Keep the homepage CTA product-oriented here, then let the docs tree handle deeper engineering and architecture reading.