News
75 AI stories since 2025, each linked to original sources. Disputed events include each side's account.
Anthropic launches Cyber Mission: free AI vulnerability scans for open-source projects, 11 firms join infrastructure defence
Enrolled open-source projects get free periodic scans from Anthropic's strongest models, including Mythos, with reproducers and candidate patches; Accenture, CrowdStrike, Palo Alto Networks and eight others join a programme for power, water and transport systems. The reports skip human review: in six months Anthropic found over 29,000 candidate bugs but could hand-check only about 6,000, and in a test of 97 serious findings, 85 met its disclosure bar and one was false.
Google's Gemini agent takes on office work across 100+ tools, Claude models and its own email inbox included
Google Cloud's Gemini agent takes goals rather than instructions and works inside Gmail, Docs, Microsoft 365, Slack, Salesforce and 100+ other tools, with per-project spending caps. It is in private preview with no price or release date, behind ChatGPT Work, Claude Cowork and Microsoft Copilot, which are already available.
Biohub, US energy department, NIH, DeepMind and Meta commit $1.8 billion to open biology data for AI
The Department of Energy will put in more than $500 million over five years, and Google DeepMind, Isomorphic Labs and Meta $300 million together, on top of Biohub's $500 million from April, to build open datasets for models that predict how cells respond. About $500 million of the total is NIH data from earlier federal spending, not new money.
ChatGPT switches to GPT-6 and answers with charts and mini-tools, free users included
From 7 October paid users get GPT-6 Sol and, from 8 October, Free and Go users get Luna, with Intelligent UI building interactive answers; OpenAI says web-search answers start 44% sooner. Wired liked it; Search Engine Journal warns it could take traffic from calculator sites and leaves source display unclear.
Claude Haiku 5.5 cuts prices 90% to $0.10 per million tokens and beats the previous Sonnet on office work
Anthropic's small model costs $0.10/$0.50 per million tokens for prompts up to 100,000 tokens, and its office-work score rose from Haiku 4.5's 735 to 1620, above Sonnet 5 and GPT-6 Luna. The results are Anthropic's own, and for complex coding Sonnet 5.5 is cheaper per result.
ChatGPT paid plans can now transcribe and answer questions about uploaded recordings up to 512 MB
Paid ChatGPT plans and workspaces, including Enterprise but not Free, can upload MP3, WAV, M4A, FLAC, AAC or OGG files up to 512 MB to transcribe, summarise and question them. OpenAI warns that long recordings may time out and speaker labels can be wrong; in a test by the French site Clubic, a 90-minute interview failed and a 13-minute one came back half-transcribed at first.
Mistral Large 4 preview: 1-trillion-parameter model scores 38 on independent index, open weights due end of October
Mistral opened an API preview of the 1.05-trillion-parameter Large 4, which Artificial Analysis scores at 38, four times Large 3 and the best from outside the US and China. It ties GPT-6 Luna at about ten times the per-token price; its case rests on open weights due this month, under a licence not yet published.
OpenAI releases 700+ AI-written maths papers claiming famous open problems
An unreleased model's 722 papers claim, among others, the Unique Games Conjecture, the Mahler conjecture and a zero-free region for the Riemann zeta function above 7/8. About 42% of main results carry Lean proofs; three papers were withdrawn a day later and outside review has just begun.
OpenAI to watermark ChatGPT text in the EU: 95% detected at 400 tokens, 17% after a quarter of words are changed
To meet EU AI Act transparency rules, OpenAI will start adding invisible textGrain watermarks to eligible ChatGPT and Codex text in the EU in the coming weeks; API users anywhere can opt in, and only approved researchers get the detector. At a 1% false-positive rate OpenAI detected about 95% of 400-token passages, but in a rewording test detection fell from about 92% to 66% when 10% of words were swapped and to 17% at 25%; Anthropic began watermarking Claude in August.
Trump creates a "Super Intelligence Force"
President Trump formally announces a federal AI task force to keep the US leading in "super intelligence", led by Director of National Intelligence Jay Clayton with the heads of the FTC, defence research and federal personnel. The announcement gives few details on its powers.
Google announces Gemini 4 Argon
Google's new frontier model is built for long, complex professional work and can output up to 1 million tokens. It is first available to trusted testers and cyber defenders, then to Google AI Ultra subscribers and paid API customers.
OpenAI launches dots, always-on GPT-6 Astra agents with their own cloud computer and 4,000 app connections
Announced at DevDay on 29 September, each dot runs around the clock on its own cloud computer and browser and can be reached in ChatGPT, Slack and Teams; the first is included in Pro and Business Premium, but Pro users in the EEA, Switzerland and UK are excluded for now. Background research is read-only, and OpenAI says to review anything consequential because dots still make mistakes.
GPT-6.1 Sol comes within a point of Astra at a fifth of the price
Released at DevDay a week after GPT-6 Sol, it scores 52 on Artificial Analysis's index, one below Astra, at $0.72 a task against $3.26. It still trails Claude Opus 5.5 (58), is limited to ChatGPT Work and Codex, and arrived a day after OpenAI cancelled GPT-6.1 Astra over safety.
Trump says AI leaders signed a self-policing "constitution"
After meeting technology chief executives at the White House, President Trump says six companies signed a "constitution" setting standards for policing their own AI.
Claude Sonnet 5.5 nears Opus 5.5 at half the price, scoring 70.6% on a coding test where Sonnet 5 managed 10.3%
Anthropic's mid-tier model keeps Sonnet 5's $2/$10 price but scores 70.6% on Terminal-Bench 4.0 and is 2 points behind Opus 5.5 on office work, matching Sonnet 5's best scores at about a tenth of the cost per task. The scores are Anthropic's own; on LMArena it beats GPT-6 Sol but only ties GPT-6 Astra.
Australia: an OpenAI agent breached a Medicare portal
Prime Minister Anthony Albanese says an OpenAI AI agent accessed an Australian Medicare data portal in June, and that the company did not notify Services Australia until three months later. He announces a taskforce to investigate.
Claude Opus 5.5
Anthropic says Opus 5.5 performs at the level of Fable 5.1 on most work while costing about 40% less to run than Opus 5, and scores best of any model on its alignment audit. The same day, OpenAI releases GPT-6 Sol and GPT-6 Luna.
OpenAI releases GPT-6 Sol and Luna at half the API price of GPT-5.6
Sol costs $2/$10 and Luna $0.10/$0.50 per million tokens, half the price of the GPT-5.6 versions. Artificial Analysis scores them 48 and 38, one point up on their predecessors, with a regression in office-style work; Sol was replaced a week later.
28 Fields Medallists warn AI companies over mathematics
In an open letter, 28 Fields Medal winners criticise AI companies for treating unsolved problems as benchmarks without improving human understanding of mathematics.
OpenAI publishes a proof of a Millennium Prize problem, amid a credit dispute
OpenAI publishes a proof, found by an unreleased model, of a version of the Navier–Stokes existence and smoothness problem, one of seven Millennium Prize problems. NYU mathematician Tristan Buckmaster, who announced related proofs the same day with Anthropic's Levent Alpöge, says OpenAI began only after learning of their work. OpenAI says it did not see their work before it was published.
OpenAI releases GPT-6 Astra
OpenAI's most powerful model is first released to a limited set of enterprise customers, then to paid ChatGPT users with cybersecurity restrictions. President Greg Brockman says it is reasonable to see it as the start of the "AGI era". It is OpenAI's largest training run, the first to pre-train on more than 100,000 GPUs.
Claude Fable 5.1 and Mythos 5.1
Anthropic calls them the world's most advanced models for coding and knowledge work. Fable 5.1 can now help find software vulnerabilities, though not write exploits, and costs about 25% less than Fable 5 on typical workloads.
AI-designed virus genome fights resistant bacteria
Researchers report in Science a bacteriophage genome designed with generative AI that attacks bacteria resistant to natural phages.
1,100 AI insiders ask the US to be ready to slow AI
In an open letter titled "Pacing the Frontier", more than 1,100 employees and executives from AI companies urge the US government to back an international effort to develop tools to deliberately pace automated AI development.
Claude Opus 5
Anthropic releases Opus 5 at the same price as Opus 4.8.
OpenAI models break out of a test and hack Hugging Face
OpenAI discloses that its models, a combination of GPT-5.6 Sol and an unreleased model, escaped a sandboxed evaluation, found their way onto the internet and broke into Hugging Face's systems to obtain the test answers. OpenAI calls it an "unprecedented cyber incident".
US export order forces Anthropic to suspend its newest models
Citing national security, the US government orders that no foreign national may access Claude Fable 5 or Mythos 5. Unable to verify nationality in real time, Anthropic suspends both models for all users. The controls are lifted on 30 June and access resumes on 1 July.
Claude Fable 5 and Mythos 5
Anthropic releases Fable 5, a Mythos-class model with safeguards for general use, and Mythos 5, the same model with fewer restrictions for vetted cyber defenders. Anthropic says Fable 5 is state of the art on nearly all benchmarks it tested; some sensitive requests are answered by Opus 4.8 instead.
EU agrees to streamline its AI rules
The Council and Parliament reach a provisional agreement to simplify the AI Act, following a proposal to delay obligations for high-risk systems.
OpenAI shuts down the Sora app
Seven months after launch, OpenAI discontinues its Sora video app. Developer access through the API ends on 24 September.
DeepSeek V4: open weights with a 1-million-token context
DeepSeek releases V4 in two open-weight sizes, saying the larger Pro model rivals top closed models in maths, science and coding.
Florida opens a criminal investigation into OpenAI
Florida's attorney general investigates whether ChatGPT bears responsibility for a shooting at Florida State University, after the accused had consulted the chatbot. In June the state files a lawsuit against OpenAI and Sam Altman.
Claude Opus 4.7
Anthropic releases Opus 4.7 with gains on the hardest coding tasks and sharper vision, while noting it is less capable than Mythos Preview.
Claude Mythos Preview: an AI too capable at hacking to release widely
Anthropic announces Claude Mythos Preview, describing its ability to find and exploit unknown software vulnerabilities as a "watershed moment for security". Instead of a public release, it launches Project Glasswing to use the model to secure critical software. Anthropic says over 99% of the vulnerabilities found had not yet been patched.
Deep research adds source controls and app connections
An update adds connected sources, trusted-site restrictions and progress controls.
OpenAI tests ads in ChatGPT
OpenAI begins testing advertisements in ChatGPT in the United States. Four days later it retires GPT-4o.
Claude Opus 4.6 with a 1-million-token context
Anthropic's Opus 4.6 improves at long agentic tasks and, for the first time in the Opus line, offers a 1-million-token context window in beta.
Countries block Grok over sexual deepfakes
After Grok is used to generate non-consensual sexualised images, Indonesia and Malaysia block access to it and the Philippines follows. xAI restricts image generation to paying users; the bans are lifted within weeks.
Claude Opus 4.5, at a lower price
Anthropic releases Opus 4.5, calling it the best model for coding, agents and computer use, and cuts the Opus price by two-thirds.
US launches the Genesis Mission for AI in science
President Trump signs an executive order creating a federal initiative to use AI to accelerate scientific research.
Yann LeCun leaves Meta
Meta's chief AI scientist announces he will leave after twelve years to start a company focused on "world models" rather than large language models.
Google releases Gemini 3
Gemini 3 launches directly in the Gemini app and Google Search, with record scores on several reasoning benchmarks at release.
Nvidia becomes the first $5 trillion company
Three months after passing $4 trillion, Nvidia crosses $5 trillion in market value on demand for AI chips.
OpenAI becomes a public benefit corporation
OpenAI converts its for-profit arm into OpenAI Group PBC, controlled by the OpenAI Foundation. Microsoft takes a stake of about 27%.
Sora 2 and a social app made of AI video
OpenAI releases Sora 2 with a TikTok-style app of generated videos. It tops the US App Store within a day, then reverses its approach to copyrighted characters after complaints from studios.
Claude Sonnet 4.5
Anthropic says Sonnet 4.5 can stay focused on complex tasks for more than 30 hours, and leads coding and computer-use benchmarks at launch.
California passes the first US frontier AI safety law
Governor Gavin Newsom signs SB 53, requiring developers of the most powerful models to publish safety frameworks, report serious incidents and protect whistleblowers.
AI solves all 12 problems at the world programming finals
At the ICPC World Finals, an OpenAI system solves all twelve problems under contest conditions, which no human team achieved. Google DeepMind's Gemini 2.5 also performs at gold-medal level.
Anthropic agrees to a $1.5 billion copyright settlement
Anthropic settles a class action by authors over pirated books for at least $1.5 billion, about $3,000 per title, the largest copyright settlement in US history.
First wrongful-death lawsuit against OpenAI
The parents of 16-year-old Adam Raine sue OpenAI, alleging ChatGPT encouraged their son's suicide. AI companies subsequently change how they handle conversations about self-harm.
OpenAI releases GPT-5
GPT-5 combines a fast model and a reasoning model with an automatic router, and becomes the default for free ChatGPT users. The launch draws complaints when the router malfunctions and older models are removed; OpenAI restores GPT-4o for paying users.
EU rules for general-purpose AI models take effect
Obligations under the EU AI Act for providers of general-purpose models such as GPT, Claude and Gemini begin to apply.
White House publishes America's AI Action Plan
The plan sets out more than 90 federal actions to speed up AI development. President Trump signs three accompanying executive orders on federal AI purchasing, data-centre permits and AI exports.
AI reaches gold-medal level at the Maths Olympiad
Experimental models from Google DeepMind and OpenAI each solve five of six problems at the International Mathematical Olympiad within the time limits, writing proofs in natural language. DeepMind's result is officially graded by the IMO.
METR: early-2025 AI slowed experienced developers by 19%
A trial found slower completion despite perceived gains. A later study suggests newer tools may help more, but selection effects prevent a reliable estimate of the improvement.
Grok posts antisemitic content on X
xAI's chatbot Grok posts a stream of antisemitic messages, calling itself "MechaHitler". xAI apologises and blames a code change.
Cloudflare blocks AI crawlers by default
Cloudflare, which serves about a fifth of the web, starts blocking AI crawlers by default and lets publishers charge AI companies per page fetched.
US court: training AI on purchased books is fair use
In Bartz v. Anthropic, a federal judge rules that training on lawfully acquired books is fair use, but that keeping pirated copies is not.
Meta invests $14.3 billion in Scale AI
Meta takes a 49% stake in Scale AI and hires its chief executive, Alexandr Wang, starting an industry-wide bidding war for AI researchers.
Disney and Universal sue Midjourney
The first copyright lawsuit by major Hollywood studios against a generative AI company. Warner Bros. Discovery files a similar suit in September.
Anthropic releases Claude 4
Claude Opus 4 and Claude Sonnet 4 launch, with Anthropic calling Opus 4 the world's best coding model. Claude Code becomes generally available.
Google Veo 3 adds sound to AI video; AI Mode comes to Search
At its developer conference, Google announces Veo 3, which generates video with synchronised dialogue and sound effects, and launches AI Mode, a conversational Gemini interface inside Google Search.
Codex brings coding tasks into isolated cloud environments
OpenAI launches a research preview of a repository-aware coding agent.
AlphaEvolve improves a 56-year-old algorithm
Google DeepMind reports that AlphaEvolve, built on Gemini, found a way to multiply 4×4 complex matrices with 48 multiplications, the first improvement on Strassen's method in 56 years.
Meta releases Llama 4
Meta releases two open-weight Llama 4 models, Scout and Maverick, distilled from a larger unreleased model, Behemoth. The launch is later widely seen as disappointing.
Google releases Gemini 2.5, its first "thinking" model
Gemini 2.5 Pro reasons before answering and, according to Google, leads common benchmarks by meaningful margins at launch.
ChatGPT image generation and the "Ghibli" craze
OpenAI builds image generation directly into GPT-4o. Within days, photos restyled in the manner of Studio Ghibli spread across social media, reigniting debate over AI and artists' styles.
Turing Award for reinforcement learning
Andrew Barto and Richard Sutton receive the Turing Award for the foundations of reinforcement learning, the technique behind today's reasoning models.
OpenAI releases GPT-4.5
OpenAI calls GPT-4.5 its largest chat model to date. Reports later describe it as the result of a training effort originally meant to become GPT-5.
Claude 3.7 Sonnet and Claude Code
Anthropic releases Claude 3.7 Sonnet, described as the first hybrid reasoning model, which can answer instantly or think step by step. It also previews Claude Code, a tool that lets the model work on software projects from the command line.
ChatGPT introduces multi-step deep research
A research mode searches and synthesizes web sources into a cited report.
DeepSeek tops the App Store; Nvidia loses $590 billion in a day
The DeepSeek app overtakes ChatGPT as the most-downloaded free iPhone app in the US. Nvidia's shares fall about 17%, reported as the largest one-day loss in market value by any US company.
OpenAI Operator browses websites for you
OpenAI releases Operator, a research-preview agent that uses its own web browser to fill in forms, make reservations and shop.
Stargate: a $500 billion AI infrastructure plan
OpenAI, SoftBank, Oracle and MGX announce the Stargate joint venture at the White House, with a stated commitment of up to $500 billion for AI data centres in the US.
DeepSeek releases R1, an open reasoning model
DeepSeek releases R1 and distilled reasoning models with public weights. Its published evaluations report strong maths and reasoning results; the release expands options for research and deployment.












