AI Watch, Friday 11 September 2026AI safety moves from the activist fringe to mainstream debate: Sam Altman says he is ready to slow frontier AI development and is testing whether such a brake would be legal, a former Anthropic researcher warns that \"AI could kill us all,\" OpenAI adds a worried researcher to its board, and Anthropic publishes 161 pages of Claude misuse, from Yemeni missiles to fake dating profiles. In Europe, the CNIL sets rules for electronic invoicing and the EU's cybersecurity agency trials Anthropic's most offensive model. In China, DeepSeek reignites the price war with its V4.1 Flash model and chipmaker Enflame jumps 188% on its stock market debut. Microsoft plans 38 gigawatts of data centres and Jensen Huang promises Nvidia 70% growth.
Par l'équipe éditoriale Masteria, sous la direction de Mathias Nizan · Publiée le à 9h09
AI safety leaves the activist fringe for mainstream debate: Sam Altman says he is ready to slow down, a former Anthropic researcher issues a chilling warning, and Anthropic publishes 161 pages on the misuse of Claude. Meanwhile, DeepSeek reignites the price war in China, Enflame soars on the market and Microsoft plans 38 gigawatts of data centres.
13 stories selected from 142 collected this morning across 38 feeds. 20 sources cited, about 8 minutes to read.
Anthropic publishes 161 pages on the misuse of Claude, from missiles to romance scams
The threat intelligence report Anthropic released on Thursday 10 September details, across 161 pages, attempts to misuse its Claude model and how the company says it blocked them. The catalogue goes far: research that could help make biological weapons, help with targeting missiles in Yemen, fake profiles run at scale on dating apps, fraud campaigns. Anthropic says it has strengthened the guardrails on its latest models to limit queries touching on sensitive biology. The document also describes "distillation" campaigns, in which Chinese competitors (Alibaba, Moonshot AI, DeepSeek) query Claude en masse to copy its behaviour and train their own models more cheaply, a practice said to have intensified in recent months. The company acknowledges a limit: it stops what it sees pass through its servers, not what is already running beyond its reach.
The 11 September edition covers 13 stories from 20 sources: 4 pour l'Europe et la France, 4 pour l'international, 2 pour la Chine et l'Asie, 1 publication de recherche et 1 brève.
Every story carries its sources. Links open the original publication.
Europe and France
4 stories
The CNIL sets rules for data protection in electronic invoicing
Since 1 September 2026, the reform of business-to-business electronic invoicing has come into force, requiring every company to issue and receive its invoices in digital form. The Commission nationale de l'informatique et des libertés (France's data protection authority) has published guidance to help companies secure the personal data that moves through these flows, in particular via the approved platforms that centralise exchanges. For finance and legal departments, the technical compliance project comes with a GDPR project that is best opened now.
Hugging Face slips a message to AI agents into its security file
The French open-model platform has added to its security.txt file, usually reserved for security researchers, a note addressed to the autonomous agents that crawl its site: no need to try to hack it, the answers to their goals are elsewhere. The wink reflects a new reality for security teams, namely that a site's visitors are no longer only human, and that agents launched by third parties now test defences around the clock.
The European cybersecurity agency trials Claude Mythos
ENISA, the European Union's cybersecurity agency, has been granted access to Mythos, Anthropic's most offensive model for cyber defence, Next.ink reports. In May, OpenAI had already opened up its GPT-5.5-Cyber to the agency. A European public institution is thus equipping itself with the same tools whose misuse Anthropic documents elsewhere, in the hope of strengthening its own defence before attackers make use of them.
Ed.ai raises €5 million for AI-assisted grading of student work
The French start-up Ed.ai has closed a €5 million funding round to develop its grading-support tool, which analyses students' work to produce more precise feedback for teachers. The project aims to lighten the marking load while keeping the teacher at the centre of assessment. French edtech finds here a concrete use case, where AI prepares the work without replacing the teacher's judgment.
AI safety becomes the central debate, and OpenAI says it is ready to slow down
Sam Altman told his teams that OpenAI was considering slowing the development of its most advanced models, and that he hoped to see other labs do the same, according to Bloomberg. At the same time, the company is trying to find out whether industry coordination to slow down might fall foul of competition law, Wired reports. The shift comes with a change of mood: Jacob Coxon, who has just left Anthropic, warns that "AI could kill us all" in the years ahead, a message picked up by Platformer, Wired and Bloomberg. OpenAI has also added to its foundation's board a researcher known for pessimism about AI risks. The existential question, long confined to a few circles in the San Francisco Bay Area, is settling at the centre of the sector's conversations.
OpenAI launches ChatGPT for Financial Services, built on GPT-6 Astra
On 10 September, OpenAI unveiled a version of ChatGPT dedicated to finance, combining built-in financial data with its latest model, GPT-6 Astra, for research, modelling and the production of client-facing documents. The tool is aimed at analysts, bankers and portfolio managers, promising to speed up market analysis and the preparation of materials. The announcement extends OpenAI's strategy of tailoring ChatGPT by sector, with the parallel launch of a "Data agent" able to connect a company's data and turn it into dashboards. Finance, a heavy consumer of document synthesis, is becoming a leading commercial target.
Microsoft plans 38 gigawatts of data centres to absorb demand
Short of capacity, Microsoft has had to turn down some AI and cloud orders, Bloomberg reveals. The group now plans to add 26 gigawatts of computing power and bring its fleet to 38 gigawatts of total capacity, the equivalent of some thirty nuclear reactors devoted to its data centres alone. The plan illustrates the current bottleneck, which owes less to chips than to electricity and available land. The AI race is now being run as much among power utilities as among model builders.
Jensen Huang promises Nvidia 70% growth and sets his sights on cybersecurity
Nvidia's chief executive, Jensen Huang, said his company would grow 70% next year, while rejecting the idea that its deals with customers are "circular," meaning that it funds those who buy its chips, according to TechCrunch. At the Goldman Sachs technology conference, he named cybersecurity as AI's next big market, after coding, citing defence systems running around the clock. Nvidia is already forming partnerships in the field, Clubic reports. The supplier of the chips that train the models is signalling where it intends to sell the next wave.
DeepSeek unveils V4.1 Flash and reignites the price war
On Thursday 10 September, DeepSeek launched its V4.1 Flash model, which it presents as more capable than its previous flagship while cutting the cost of inference (the cost of running a model once it is trained) and gaining speed. The model rests on a new architecture called "Causal-Encoder-Decoder" and on 552 billion parameters, with a "mixture of experts" system (only a fraction of the network activates on each query, which limits computation), whereas a conventional model draws on the whole network. DeepSeek claims V4.1 Flash beats Moonshot AI's Kimi K3 on cybersecurity and coding benchmarks. The announcement fits China's strategy of aggressive pricing and comparable performance, which is pushing down global AI prices.
Chinese chipmaker Enflame jumps 188% on its Shanghai stock market debut
Enflame Technology, one of Nvidia's main Chinese rivals, saw its share price climb 188% on Friday 11 September in its first trading session in Shanghai. The stock opened at 410 yuan, against an offer price of 142.18 yuan, taking the valuation of this Tencent-backed AI chip designer to 176.4 billion yuan, around $25 billion. Enflame was the last of China's four big AI chip champions to go public. The enthusiastic investor reception confirms the domestic market's appetite for local alternatives to Nvidia, as US export controls push China to build its own supply chain.
Two public corpora finally make it possible to measure, without guessing, what a model has actually "memorised."
A team takes on a question that plagues lawsuits over training data: when a model predicts a sentence with suspicious ease, can you conclude that it was in its training corpus? Almost every test published so far had to guess which sentences were present. The authors remove the guesswork by relying on two model families, OLMo-2 and Pythia, whose corpora are public and indexed, which gives the exact number of times any sentence appears. Their conclusion tempers expectations: the membership signal is detectable only where duplication blurs with other factors, which weakens the methods meant to prove that a text was used for training. This should shed light on ongoing disputes, including the accusation levelled this week against OpenAI over the origin of its mathematical data.
Claude Mythos has found 26,000 flaws in five months, fewer than 1% have been fixed. Five months after its launch, Anthropic's vulnerability-hunting tool shows impressive numbers, but analysts point out that detection is worth nothing without repair, when more than 25,000 flaws remain open. Numerama
Moonshot AI, maker of Kimi, is studying a dual listing in Hong Kong and Shanghai. The Chinese lab, developer of the Kimi K3 model, is considering a listing on both exchanges to raise capital and gain visibility, after negotiating a Hong Kong entry in July, according to two sources cited by the South China Morning Post.
Sam Altman tells his teams that OpenAI is ready to slow frontier AI and is testing the legality of such a brake, while fear of extinction floods the news feeds: the reflex that marks a serious organisation is to build its risk register on the misuse already documented, like Anthropic's 161 pages, and to measure the share of flaws fixed rather than the number found, when Claude Mythos finds 26,000 in five months and fewer than 260 are repaired
Sam Altman told his teams that OpenAI is ready to slow the development of frontier AI, and the company is already trying to find out whether such a brake would be legal under competition law. In a matter of days, fear of extinction moved from the rationalist circles of the San Francisco Bay Area into mainstream debate: a former Anthropic researcher, Jacob Coxon, warns that \"AI could kill us all,\" OpenAI adds a worried researcher to its board, and Wired and Bloomberg devote their podcasts to it. The risk of 2030 fills the news feeds. The risks of this quarter, meanwhile, are already catalogued, dated, and almost never addressed. Anthropic has just published the inventory: 161 pages of Claude misuse, from Yemeni missiles to fake profiles on dating apps. The same day, Numerama recalled a figure that had gone unmentioned. In five months, Anthropic's vulnerability-hunting tool, Claude Mythos, has found 26,000 of them in real software. Fewer than 260 have been fixed. More than 25,000 remain open. The machine finds, no one repairs. That is the real gap, and it is not speculative. An organisation's seriousness shows in its risk register: what is on it, what has been fixed, how long it took. Take every AI system running in your organisation. Test it against the misuse the Anthropic report documents: hidden instruction injection, credential theft delegated to an agent, data exfiltration. Then measure what counts: the share fixed and the time to fix, not the number of flaws found. Name who closes each hole, the way you name someone responsible for payroll. The labs debate the end of the world and ask the law whether they are allowed to slow down. Meanwhile, a workflow running in your organisation is waiting for a fix no one has planned. The distant threat holds attention. The near threat awaits a decision. It is yours to make. The Masteria editorial team.
Sam Altman told his teams that OpenAI is ready to slow the development of frontier AI, and the company is already trying to find out whether such a brake would be legal under competition law. In a matter of days, fear of extinction moved from the rationalist circles of the San Francisco Bay Area into mainstream debate: a former Anthropic researcher, Jacob Coxon, warns that "AI could kill us all," OpenAI adds a worried researcher to its board, and Wired and Bloomberg devote their podcasts to it. The risk of 2030 fills the news feeds. The risks of this quarter, meanwhile, are already catalogued, dated, and almost never addressed.
Anthropic has just published the inventory: 161 pages of Claude misuse, from Yemeni missiles to fake profiles on dating apps. The same day, Numerama recalled a figure that had gone unmentioned. In five months, Anthropic's vulnerability-hunting tool, Claude Mythos, has found 26,000 of them in real software. Fewer than 260 have been fixed. More than 25,000 remain open. The machine finds, no one repairs. That is the real gap, and it is not speculative.
An organisation's seriousness shows in its risk register: what is on it, what has been fixed, how long it took. Take every AI system running in your organisation. Test it against the misuse the Anthropic report documents: hidden instruction injection, credential theft delegated to an agent, data exfiltration. Then measure what counts: the share fixed and the time to fix, not the number of flaws found. Name who closes each hole, the way you name someone responsible for payroll.
The labs debate the end of the world and ask the law whether they are allowed to slow down. Meanwhile, a workflow running in your organisation is waiting for a fix no one has planned. The distant threat holds attention. The near threat awaits a decision. It is yours to make. The Masteria editorial team.
The Masteria editorial team
Under the direction of Mathias Nizan
Il forme les équipes dirigeantes et techniques à l'IA générative depuis 2022.
38 feeds were reviewed on the morning of 11 September, 142 stories collected, 13 selected, each linked to its source. The analysis is written by the editorial team and published with the edition.
Sources du jourFrandroidLe MondeTechCrunchCNILNumeramaNext.inkMaddynessBloombergWiredSiècle Digital