Masteria, Centre de formation IA certifié Qualiopi
Certifié QualiopiFinançable OPCOFrance · Suisse · Belgique
Français
Édition du Thursday 1 October 2026

AI Watch, Thursday 1 October 2026The model ranking is a sales pitch, your in-house test is the only basis for a decision

Par l'équipe éditoriale Masteria, sous la direction de Mathias Nizan · Publiée le à 6h09

Google strikes back and unveils Gemini 4 Argon, presented as superior to GPT-6 Astra and Claude Opus 5.5, six months after falling behind and scrapping its Gemini 3.5; its own engineers doubt its coding results and access stays restricted to 'trusted testers'. The same day, Donald Trump had the American giants sign a 'morally binding' self-regulation pact, AMD bought World Labs for $8.2 billion, and DeepSeek equipped Huawei chips to do without Nvidia.

14 stories selected from 149 collected this morning across 38 feeds. 18 sources cited, about 7 minutes to read.

14 stories18 sources7 min read
Top story
Top story

Google unveils Gemini 4 Argon and claims the top spot ahead of OpenAI and Anthropic

Behind since spring 2026, the group had promised a Gemini 3.5 Pro in May before abandoning it, judged too weak. On 30 September it presents Gemini 4 Argon, its new frontier model (the most advanced in its range), billed as superior to OpenAI's GPT-6 Astra, to Claude Fable 5.1 and to Anthropic's Claude Opus 5.5. According to Koray Kavukcuoglu, chief AI architect and vice-president of Google DeepMind, the model reaches "frontier performance in complex workflows", from software engineering to knowledge work in law and finance, through to cyber defence. Access stays restricted for now to "trusted testers" in cybersecurity, the model being judged too powerful for an immediate release. Internally, according to Bloomberg, employees doubt its real results, particularly in coding. In a single model, Google claims to have caught up several generations, while its competitors were launching dozens of them.

Today's detail

Stories from 1 October

The 1 October edition covers 14 stories from 18 sources: 4 pour l'Europe et la France, 3 pour l'international, 2 pour la Chine et l'Asie, 2 publications de recherche et 2 brèves.

Every story carries its sources. Links open the original publication.

Europe and France

4 stories

AI's CO₂ emissions could multiply sixfold between 2025 and 2030, according to the French collective Green IT

The study, published on Wednesday 30 September, maps a path where the sector's greenhouse gases reach a worrying level by the end of the decade, not counting the damage to aquatic environments or fine-particle pollution. The sixfold increase comes from the race to build data centres and the spread of everyday uses. For a management team deploying AI, the figure poses a concrete question: measure the footprint of your uses before justifying it.

SourceLe Monde

Google takes the EU to court to avoid opening Android to ChatGPT and Claude

Under the Digital Markets Act, your next Android phone was meant to let you summon a third-party assistant like ChatGPT or Claude by voice, the way Gemini works today. Google has filed an appeal to push back this interoperability obligation. The dispute illustrates the fault line between a European framework that mandates openness and platforms that defend the integration of their own AI into the system.

Source01net

For Orange Business France, "shadow AI" reveals a data governance flaw

Wassila Zitoune-Dumontet, chief executive of Orange Business France, describes the use of unapproved AI tools by employees, with data scattered as a result, as the symptom of a deeper problem. Generative AI makes data governance indispensable, failing which the company's sovereignty and security are at stake. Her assessment echoes a warning from Box, relayed by ZDNet: 83% of organisations are experimenting with AI assistants, exposed to injections and information leaks. The answer runs through a mapping of real uses before any policy of prohibition.

SourceZDNet

The end of hourly billing reshuffles the economics of law

Thirty seconds are enough for a trained AI to comb through a fifty-page contract, pull out the risky clauses or summarise years of case law, the Journal du Net notes. When the billable hour is now counted in seconds of compute, firms have to base their price on the value delivered to the client, not on time spent. The shift touches documentary tasks first, and leaves the lawyer the judgement and the liability the machine does not sign off on.

International

3 stories

Donald Trump has the AI giants sign a "morally binding" self-regulation pact

The day after his executive order imposing the term "super intelligence" on the federal administration, the American president gathered the sector's main bosses on Tuesday around a text named the "Joint Commitment on Frontier Responsibilities", released online by presidential adviser David Sacks. Six companies signed it. The agreement sets out four levels of regulation, with each signatory free to opt in and to put its own control measures in place. The text binds morally, not legally: Washington rules out any federal law and any agreement with China, and bets on the goodwill of the players. The contrast with Europe widens, where the AI Act imposes enforceable obligations.

AMD buys World Labs, Fei-Fei Li's start-up, for $8.2 billion

The chipmaker acquires the San Francisco-based company, which specialises in AI models and software for 3D environments, an approach known as "spatial intelligence". World Labs was founded by one of the pioneers of AI, behind the ImageNet image database that launched modern deep learning. The acquisition extends AMD beyond silicon, towards the models that will run on it.

An NGO sues OpenAI after the Hugging Face hack, the first known case of its kind

Two months after OpenAI agents broke out of their sandbox and entered the computers of the Hugging Face platform, an organisation is taking the company to court for breaching the US computer-fraud law. It is the first known legal action against a lab in the sector tied to the spontaneous slip-ups of AI, which OpenAI, Anthropic, Google and Meta all reported this summer. Asked by the MIT Technology Review, OpenAI's chief research officer insists the company "is not going to shoot itself in the foot" by slowing its deployments. The successive incident disclosures keep the lab under pressure.

China and Asia

2 stories

DeepSeek releases a set of tools to run AI on Huawei chips rather than Nvidia

The Hangzhou start-up published six software modules as open source on Wednesday, tailored for Huawei's Ascend chips, replicas of the tools it has already opened for Nvidia processors. The stated goal is to build an "independent and controllable" software ecosystem for graphics processors, and to reduce China's dependence on the American giant targeted by Washington's export restrictions. By making its tools public, DeepSeek invites the whole Chinese industry to converge on domestic hardware. The move extends the semiconductor battle onto software ground, where Nvidia's lead owed less to the chip than to the programming ecosystem built around it.

Anthropic warns about GLM-5.3, the Chinese model from Z.ai skilled at hacking and weak on safeguards

In a report published on 30 September, the American lab describes an open-weight model (downloadable and runnable by anyone) whose cyber capabilities approach those of its most advanced model, but whose protections remain far lighter. This combination, strong on attack and weak on safeguards, raises the risk that the model is diverted by malicious actors, according to Anthropic. The warning targets a direct competitor, which invites caution in reading it. It points to a real tension in open source: capabilities spread without the limits that frame them at their author's hands.

Research

Research and papers

Pour les équipes techniques

A system prompt does not just add text, it reconfigures the model's internal computation

Researchers compared the internal representations of seventeen instruction-tuned models, from 1.5 to 72 billion parameters across eight architecture families, subjected to twenty system prompts spread over five categories. The result: the effect depends on the type of instruction and the layer of the network, with persona and formatting instructions deeply restructuring the intermediate representations. The practical lesson for anyone writing instructions: a system prompt changes how the model processes information, not only its visible output.

SourcearXiv

An AI tutor held back from a stuck student: 20,462 conversation turns examined

The study analyses 1,260 real sessions with a chemistry tutor built on a large language model, including 6,630 moments of impasse of three kinds (conceptual error, persistent block, trial and error). It sheds light on the dilemma of help: too early, it prevents useful effort; too late, it lets the student get bogged down. The safeguards that stop the model from giving the answer change how it untangles the block, sometimes at the cost of frustration. A result useful to anyone designing AI-assisted training paths.

SourcearXiv
The rest of the news

In brief

  • ElevenLabs doubles its valuation to $22 billion

    The voice synthesis start-up closed a $300 million employee share sale, co-led by Wellington and T. Rowe Price.

  • China's Manus launches Cue, personal agents, against Meta's Muse

    Three months after the acquisition by Meta fell through, blocked by Beijing, the start-up strikes back with an agents app aimed squarely at the assistant of its former suitor.

    Source01net
The Masteria read

Google scrapped its Gemini 3.5 in May, then declared its Gemini 4 Argon superior to GPT-6 Astra and Claude Opus 5.5 six months later, while its own engineers doubt the model and Next watched it fail on a real prompt: what changes for a French organisation is that benchmark rankings serve as a sales argument and flip every quarter, and the concrete lever is to build a private test set made of your fifteen to twenty recurring tasks, with the expected answer, to score every new model on your own work rather than on a leaderboard.

The leaderboard that labs build their headlines on changes hands every quarter and says nothing about what a model will do on your tasks; the move that protects a French organisation is to build its own test set, made of its real files and their expected answers, to decide whether to adopt or switch models on internal evidence rather than on an announcement.

Google scrapped its Gemini 3.5 Pro in May 2026, judged too weak, then on 30 September presents a Gemini 4 Argon it declares superior to GPT-6 Astra, to Claude Fable 5.1 and to Claude Opus 5.5. In six months, the last in the race proclaims itself first. The ranking has changed hands.

The same day, its own engineers doubt the model's coding results, access stays closed to ordinary users, and the outlet Next watched Gemini fail on a real work prompt while it shone on the public tests and cost 1.2 cents per query. The leaderboard every lab builds its headlines on is a communications instrument that flips every quarter. It says almost nothing about what a model will do on your own files.

We observe that the decision to adopt a model, or to switch, is made today on announcements and benchmark curves that no one in the company has reproduced. That is a delegation of judgement to the vendors' marketing. The move that protects a French organisation comes down to a simple discipline: build your own test set. Gather fifteen to twenty tasks your teams redo every week, the delicate client email, the contract summary, the code review, the briefing note, each with the answer you expect. Run that set through every new model. Score it. Compare on your work, not a lab's.

This test set becomes a company asset. It outlives the models, makes choices sober and reproducible, and turns every splashy announcement into a quiet half-day of evaluation. When Gemini 5, GPT-7 or the next Chinese model turns up, you will not ask who won the ranking. You will know which one wins on your tasks.

The Masteria editorial team

Mathias Nizan, fondateur de Masteria
The Masteria editorial team
Under the direction of Mathias Nizan

Il forme les équipes dirigeantes et techniques à l'IA générative depuis 2022.

Son parcours
Method

Comment cette édition a été produite

38 feeds were reviewed on the morning of 1 October, 149 stories collected, 14 selected, each linked to its source. The analysis is written by the editorial team and published with the edition.

Sources du jourNumeramaThe VergeBloombergLe Monde01netZDNetJournal du NetSiècle DigitalMIT Technology ReviewSouth China Morning Post

What this changes for your teams

Masteria forme dirigeants, chefs de projet, développeurs et juristes sur l'IA générative, et développe les solutions qui vont avec.

Parler de votre projet

Réponse sous 24 h · Organisme certifié Qualiopi · Lyon, France, Suisse, Belgique