News World

Artificial Intelligence
Vol. I · Archive 10 entries

Releases

Models, APIs and products as they shipped, with what they do, what they do not do, and the link to the original announcement.

Releases

Issue 002 · A newsroom with no journalists

GPT-6 Astra

$10 per million input tokens, $50 per million output tokens. Rolling out across ChatGPT Plus, Pro, Business, Enterprise and the OpenAI API.

What it's for Programming, cybersecurity, scientific research, computer control. It is the first model to cross the "critical" threshold of OpenAI's preparedness framework in cybersecurity, meaning it can attack as well as defend. A fast mode doubles the speed, and doubles the price.

The limit The rollout is not finished, not everyone has it yet. And it is a model you cannot host yourself: output runs at $50 per million tokens, where Google's Flash range sits at $3.75.

openai.com/index/gpt-6-astra

Gemini 3.8 Flash (and Cyber)

Google's latest Flash model, aimed at agents, reasoning and cybersecurity.

What it's for Running agents and reasoning at the lowest cost in the Google lineup. The Cyber variant is reserved for defence teams invited by Google.

The limit Flash Cyber is not publicly available, it is reserved for defence teams. The price appears nowhere on the announcement page.

blog.google/innovation-and-ai/…

DeepSeek-V4-Flash-Vision-Exp

184,542 downloads, 695 likes on Hugging Face. MIT licence.

What it's for Reading text and images, for free, with weights you host yourself. Useful for powering an agent that needs to understand screenshots or scanned documents.

The limit Experimental build (it says so in the name). No technical paper spotted.

huggingface.co/deepseek-ai/…

SlotStream

227 points and 114 comments on Hacker News, number one on Show HN on 1 September 2026. Open source, MIT licence.

What it's for Running a large model (Qwen3.8-Flash-Next, 105 GB) on a Mac with 48 GB of memory, at roughly 12 tokens per second, without sending anything outside.

The limit macOS only, with enough unified memory. 12 tokens per second is fine for asking questions, not for processing volume.

github.com/carloslfu/slotstream

OpenMontage

56,292 stars, 7,065 forks on GitHub. AGPL-3.0 licence (strong copyleft: if you expose a service built on it over a network, you must publish your modifications). No official release published.

What it's for You describe the video you want, and the tool chains search, script, images, voice, music, editing and rendering on its own. 12 pipelines available, including Clip Factory, which cuts one long source into a batch of ranked short clips. The project's demos are priced between $1.33 and $5 per video, with a default cap at $10.

The limit Heavy install (Python, Node, Remotion). No ready-to-run release.

Watch out for the lookalike. The repo OpenMontage-app/OpenMontage (1 star, created 27 August 2026) copies the real project's description word for word without being a fork, and replaces the AGPL with an MIT licence. Do not confuse the two.

github.com/calesthio/OpenMontage

ScrapeGraphAI

30,601 stars, 3,057 forks on GitHub. Version 2.2.2, MIT licence. Python 3.12 minimum.

What it's for You give it a page address and a sentence describing what you want out of it. A model reads the page and hands back structured data. The library is free, the model calls are on you. The hosted service starts at $20 a month.

The limit Does not read RSS feeds or X. JavaScript rendering is on you (Playwright to install). Supported sources are websites and local documents.

github.com/ScrapeGraphAI/…

Kimi K3

2,552,594 downloads, 11,209 likes on Hugging Face. 8,710 stars on GitHub. House "Kimi K3" licence, open weights with conditions (above $20 million in revenue from reselling inference, you need an agreement with Moonshot AI).

What it's for Understanding text and images, and swallowing a 1,048,576-token window at once, roughly three quarters of a million words. Useful for summarising very large documents or feeding an agent that needs to read a lot. Text now reads through llama.cpp.

The limit It is not meant to run on your machine. The smallest light build weighs 466 GB, the others go up to 594, 861 and 1,509 GB. Vision is not yet supported by llama.cpp, only text is.

huggingface.co/moonshotai/Kimi-K3

KIMI K3, THE SMALLEST FILE BY COMPRESSION LEVELUD-Q1_0 build466 GBIQ1_S build594 GBQ2_K_XL build861 GBQ4_K_XL build1,509 GBA developer's Mac48 GB of memory

Issue 000 · Your agent succeeds one time in three

Qwen3.8-Flash-Next

What it's for the small tasks you run a thousand times a day, where every second and every cent counts, and reading images alongside text.

The counters 208,000 downloads, 4,580 likes, plus 431,000 downloads of the light build.

The limit the licence is the vendor's own, not Apache: read it before you make money with it. No image generation.

huggingface.co/Qwen/Qwen3.8-Flash-Next, and unsloth/Qwen3.8-Flash-Next-GGUF for the build that runs on your own machine, with nothing leaving it.

GLM-5.3-Flash, alias Ox Alpha

The story of the week. A model with no vendor name appeared on OpenRouter, free, under the handle Ox Alpha, and started beating the best on the leaderboards. Public detective work, tokenizer analysis, online betting. Then the Chinese lab Z.ai confirmed it itself: this was its GLM-5.3-Flash, put there as a full-scale live test to collect real usage before the announcement. The weights are now downloadable.

What it's for carrying a long project without losing the thread, with text and images mixed. Sparse architecture, 320 billion parameters of which only 18 work on each word: that is what makes it cheap to run.

The counters 441,348 downloads last month, 1,930 likes, MIT licence, 300,000 words of context.

The limit nobody has tested it independently since it was unmasked, and it is far too heavy to run on your machine: you go through a hosted service.

huggingface.co/zai-org/GLM-5.3-Flash, and Z.ai's own announcement lifting the pseudonym.

MiniMax-H3, video faster than real time

What it's for turning a sentence or a picture into video, without the wait.

The counters 5.5 million downloads of the open weights.

The limit the model is free but heavy, and the service that streams it live is paid per use.

huggingface.co/MiniMaxAI/MiniMax-H3

Back to the archive

The archive reproduces each issue as it went out, without rewriting it. A repo's stars, prices and last-commit dates are those of the day of the check, not today's: follow the link for the current state.

News World AI

Geneva, Switzerland. Write to hello@newsworldai.xyz.