AI Stream Online mascot AI STREAM ONLINE The World of AI. Live.

Methodology

Last updated: September 2026

AI Stream Online is built entirely on automated pipelines and openly available data. This page explains how every number and story on the site is produced, where it comes from, and — just as important — what it does not capture. We'd rather tell you a metric is incomplete than let it look more authoritative than it is.

1. News aggregation

Articles are pulled from public RSS/Atom feeds on an automated schedule that runs every 3 hours. The current sources are the AI sections of TechCrunch, Ars Technica, MIT Technology Review and The Verge; the AI-focused outlets AI News, The Decoder, AI Business and IEEE Spectrum; the NVIDIA, Google DeepMind and MIT News research blogs; and Rest of World and the South China Morning Post for coverage from outside the usual US/EU sources. Each run merges newly fetched articles into a rolling archive (capped at 100 current items), deduplicated by article URL. We link to and credit the original source for every story; we never republish full article text.

Feeds that are not AI-only — the research blogs and the general-interest outlets — pass through a keyword relevance check first, so a chip earnings report or a phone review doesn't land here as AI news. That check is deliberately broad, so the occasional loosely-related story gets through.

2. Story clustering

When two articles from different feeds share at least one tagged company and have substantially similar titles on the same day, they're merged into a single story with an "Also covered by" credit to the other outlet(s), instead of appearing as unrelated duplicates. Company overlap alone isn't enough on its own to trigger a merge, so two unrelated stories that happen to mention the same company aren't grouped together.

3. Per-article intelligence tagging

Every article is automatically tagged with an impact level, an automated sentiment score, and mentioned companies, countries, and technologies (labeled "Related companies & topics" on each card — they're extracted name matches, not an analysis of who is actually affected). This is done with a keyword/pattern-based heuristic engine, not a large language model. Country tagging only fires when an article's text explicitly names a country, so per-country news coverage is sparser than per-company coverage — a story about a company isn't necessarily about the country it's headquartered in.

Impact level: High and Medium are assigned only when specific keyword signals are found. When neither matches, the article shows no impact badge at all — a "Not assessed" state, not an assessed "Low."

Automated sentiment (Positive/Negative/Neutral) is a keyword count, not a factual verdict or a prediction of real-world impact — it says the engine spotted more positive- or negative-coded words in the text, nothing more.

4. AI companies tracked by country

Company counts shown per country come from Wikidata's public SPARQL endpoint — a free, live, structured, community-edited database. Companies are identified by their AI-industry tag and grouped by headquarters country.

Known limitation, stated plainly: Wikidata is edited by volunteers, so coverage is uneven and skews toward companies that English-speaking contributors happened to document. It is not a census. As one concrete example, China's real AI industry is vastly larger than the handful of companies currently tagged this way on Wikidata. We label this metric "AI companies tracked on Wikidata," never "total AI companies," for exactly this reason. Refreshed automatically every 3 hours.

A country gets its own page (and each of its companies a page) once it has at least 3 tracked companies, which filters out a small amount of data-quality noise below that floor. Entities with no usable name on Wikidata are also excluded from every count and list. This threshold is data-driven and can shift as Wikidata's coverage changes — it is not a fixed, curated list.

5. Explore AI by Country

This is an unranked, unordered-by-us grid of every country with at least 3 tracked companies (41 as of this writing), sorted only by that real Wikidata count — never described as a measure of "AI leadership." Each entry links to that country's full page (recent tagged stories, the same company list, always live).

The shaded world map above the list is the same dataset, not a second metric: each tracked country is placed into one of six real, data-derived count ranges and colored from dark brown to bright gold accordingly, with a legend showing which color maps to which range. It's a visual index into the exact numbers already in the list below — hovering or clicking a country shows its precise count. The base map itself is a public-domain map from Wikimedia Commons, simplified for file size; country shapes are for visual reference only and carry no data of their own beyond the coloring we apply.

6. GitHub activity (open-source AI sample)

The GitHub numbers (stars, open issues) are fetched from GitHub's official public API for a fixed list of widely-used open-source AI projects — a sample of open-source AI activity, not a measure of AI activity overall, and we say so next to the numbers. We show the last successfully fetched value and how long ago it was fetched, rather than a continuously ticking extrapolated estimate between refreshes. If a fetch fails, the whole widget hides rather than show a stale or zeroed-out number as current.

7. Videos and podcasts

Featured videos are selected via the YouTube Data API across several distinct AI sub-topics, taking the top result from each by view count so one viral topic can't crowd out the rest. Podcast episodes are pulled directly from each show's own public RSS feed.

8. Reader poll ("How Will This Affect People?")

Each story shows two independent things, kept visually separate: an automated sentiment read (see section 3) and a reader poll asking "How will this affect people?" with Positive / Negative / Not Sure as answers.

The poll is backed by a real shared counter with basic per-IP rate limiting and input validation on the server. When that backend isn't attached, the widget shows "Your reaction: [your answer]" instead of a fabricated or purely-local "community" figure — your own answer is never presented as if it represents other readers. Honest limitation: duplicate-vote prevention is per-browser only — nothing stops the same person from answering again in a different browser or after clearing site data. This is a lightweight reaction poll, not a scientifically controlled survey.

9. Update frequency

News, stock quotes, videos, GitHub activity, podcasts, and AI company counts all refresh automatically every 3 hours via a scheduled job. The Top AI Platforms / Top AI Voices lists are editorial and updated periodically by hand, not on a fixed schedule.

10. What we deliberately don't do

We don't fabricate metrics to fill a gap. Where we don't have a reliable free data source — for example, a genuine live count of "how many people use AI worldwide," or a comprehensive count of AI companies from trademark or social-media data — we say so and leave the metric out, rather than presenting an invented number as fact.

11. Questions or corrections

If you spot a data error or have a source suggestion, contact us at contact@aistreamonline.com.

AI STREAM ONLINE

© 2026 aistreamonline.com — All rights reserved

Privacy Policy Terms of Use Methodology