Top frontier AI models
Ranked by Intelligence Index – a composite quality score (0–100) from Artificial Analysis measuring reasoning, coding, and knowledge tasks. Higher is better. Click any column header to sort.
# Model name + open-source / modality badges Company who makes it Artificial Analysis Intelligence Index blended quality score, 0–100 SciCode scientific coding, 0–100% Humanity's Last Exam hard academic questions, 0–100% AA-Omniscience accuracy factual questions answered correctly, 0–100% GDPval-AA v2 agentic real-world work tasks, 0–100% API price CHF / 1M tokens Speed tokens / second Context window k tokens (1k ≈ 750 words)
1 Claude Fable 5.1 (max with fallback)Vision Anthropic
53.4
63% 59% 67% 63% CHF 5.87/M 68 t/s 1000k
2 Claude Fable 5.1 (xhigh with fallback)Vision Anthropic
53.2
61% 59% 66% 62% CHF 5.87/M 58 t/s 1000k
3 GPT-6 Astra (max)Vision OpenAI
52.8
56% 55% 63% 54% CHF 6.30/M 53 t/s 1000k
4 GPT-6 Astra (xhigh)Vision OpenAI
52.5
56% 55% 62% 53% CHF 6.30/M 53 t/s 1000k
5 Claude Fable 5.1 (high with fallback)Vision Anthropic
51.2
59% 56% 65% 57% CHF 5.87/M 58 t/s 1000k
6 GPT-6 Astra (high)Vision OpenAI
51.0
55% 53% 61% 51% CHF 6.30/M 52 t/s 1000k
7 Claude Opus 5 (max)Vision Anthropic
50.7
56% 55% 61% 62% CHF 3.15/M 53 t/s 1000k
8 GPT-6 Astra (medium)Vision OpenAI
49.7
54% 53% 61% 50% CHF 6.30/M 53 t/s 1000k
9 Claude Opus 5 (xhigh)Vision Anthropic
49.7
56% 54% 60% 60% CHF 3.15/M 54 t/s 1000k
10 Claude Fable 5.1 (medium with fallback)Vision Anthropic
49.1
56% 54% 63% 54% CHF 5.87/M 54 t/s 1000k
11 Claude Opus 5 (high)Vision Anthropic
48.2
55% 53% 59% 56% CHF 3.15/M 54 t/s 1000k
12 Muse Spark 1.3 (max)VisionVideo Meta
48.2
59% 49% 44% 60% CHF 0.64/M 255 t/s 1000k
13 GPT-5.6 Sol (max)Vision OpenAI
47.1
57% 49% 59% 56% CHF 2.52/M 73 t/s 1000k
14 Claude Fable 5.1 (low with fallback)Vision Anthropic
47.0
57% 49% 60% 50% CHF 5.87/M 54 t/s 1000k
15 GPT-6 Astra (low)Vision OpenAI
46.0
54% 49% 60% 46% CHF 6.30/M 51 t/s 1000k
16 Qwen3.8 Max (0902)VisionVideo Alibaba
45.4
52% 43% 32% 59% CHF 0.96/M 40 t/s 984k
17 Muse Spark 1.3 (xhigh)VisionVideo Meta
45.2
60% 47% 42% 58% CHF 0.64/M 241 t/s 1000k
18 Claude Opus 5 (medium)Vision Anthropic
45.1
52% 51% 57% 51% CHF 3.15/M 53 t/s 1000k
19 GLM-5.3 (max)Open Z AI
44.9
59% 42% 34% 57% CHF 0.74/M 67 t/s 1000k
20 Grok 4.6 (high)Vision SpaceXAI
44.4
56% 43% 48% 57% CHF 1.10/M 64 t/s 500k
How to read: The Artificial Analysis Intelligence Index (0–100) is a single blended quality score. The four columns immediately after it break that down: SciCode = writing working scientific analysis code (0–100%); Humanity's Last Exam = 2,158 frontier academic questions across maths, humanities and natural sciences (0–100%); AA-Omniscience accuracy = the share of 6,000 factual questions across 42 topics answered correctly (0–100%); GDPval-AA v2 = 220 real-world professional tasks across 44 occupations and 9 industries, scored as an Elo rating from blind pairwise comparisons and normalised to 0–100% as (Elo − 500) / 2000, with a human professional baseline of 1,000 Elo (50%). It is a relative win-rate score, not the share of tasks completed correctly. Accuracy counts correct answers only and says nothing about whether a model admits ignorance or invents an answer – Artificial Analysis publishes a separate hallucination rate and a combined Omniscience Index for that. Open = model weights are publicly downloadable. Vision / Video = input modalities beyond text. Open-weights status comes from the Artificial Analysis dataset; modalities were confirmed per model in the provider's own API documentation (Anthropic, OpenAI, Meta, Alibaba Cloud Model Studio, xAI). A missing badge means "not confirmed", not "not supported" – GLM-5.3 carries no modality badge because Z AI had not published confirmed input modalities at the time of writing. No model in this top 20 has confirmed audio input. Price = CHF per 1M tokens, converted at USD 1 = CHF 0.82 (rate of 16 September 2026) from Artificial Analysis's blended rate, which assumes a 7:2:1 ratio of cache-hit : input : output tokens. It is a per-token rate, so it is the same for every reasoning-effort setting of a model – what changes between settings is how many tokens the model spends, not the rate. Speed ≈ words/sec. Context = max text read at once. "not reported" means the value is not published for that model; such cells always sort to the bottom. All 20 models in this table have a published score for all four benchmarks. Suffixes such as (max), (high), (medium), (low) are reasoning-effort settings of the same model, scored separately; "(with fallback)" means the vendor routes overflow requests to a fallback model. Data: Artificial Analysis (SciCode, Humanity's Last Exam, AA-Omniscience) – verify before decisions.
What's new this week
Announcements from AI companies – items older than 72 hours are date-stamped. Click "Source" to read the original.
New AI models released
Google – Gemini 3.8 Live
Two new voice models for real-time spoken conversation: Gemini 3.8 Live for cheaper everyday dialogue, and Gemini 3.8 Live Extended Thinking for multi-step reasoning. Both handle 97 languages with automatic detection, can look at images mid-conversation, and watermark all generated audio with SynthID. Available now in the Gemini API and AI Studio. Promising for multilingual interview or helpline prototypes – but settle data residency before recording any participant. Zwei neue Sprachmodelle für Echtzeit-Gespräche: Gemini 3.8 Live für kostengünstigere Alltagsdialoge und Gemini 3.8 Live Extended Thinking für mehrstufiges Schlussfolgern. Beide beherrschen 97 Sprachen mit automatischer Erkennung, können mitten im Gespräch Bilder betrachten und versehen alle erzeugten Audios mit SynthID-Wasserzeichen. Ab jetzt in der Gemini API und in AI Studio verfügbar. Interessant für mehrsprachige Interview- oder Helpline-Prototypen – klären Sie aber den Datenstandort, bevor Sie Teilnehmende aufzeichnen. Deux nouveaux modèles vocaux pour la conversation en temps réel : Gemini 3.8 Live pour le dialogue courant à moindre coût, et Gemini 3.8 Live Extended Thinking pour le raisonnement en plusieurs étapes. Les deux gèrent 97 langues avec détection automatique, peuvent regarder des images en cours de conversation et filigranent tout audio généré avec SynthID. Disponibles dès maintenant dans l'API Gemini et AI Studio. Prometteur pour des prototypes d'entretien ou de ligne d'assistance multilingues – mais réglez la question de la résidence des données avant d'enregistrer des participants. Due nuovi modelli vocali per la conversazione in tempo reale: Gemini 3.8 Live per il dialogo quotidiano a costo ridotto e Gemini 3.8 Live Extended Thinking per il ragionamento in più passaggi. Entrambi gestiscono 97 lingue con rilevamento automatico, possono guardare immagini durante la conversazione e applicano la filigrana SynthID a tutto l'audio generato. Disponibili da ora nell'API Gemini e in AI Studio. Interessanti per prototipi multilingue di intervista o helpline – ma chiarite la residenza dei dati prima di registrare i partecipanti. 两款面向实时语音对话的新模型:Gemini 3.8 Live 面向成本更低的日常对话,Gemini 3.8 Live Extended Thinking 面向多步推理。两者均支持 97 种语言并可自动识别,能在对话中查看图像,并为所有生成音频加上 SynthID 水印。现已在 Gemini API 和 AI Studio 中提供。可用于多语种访谈或热线原型——但在录制参与者之前须先确认数据存放地。
Source →
Sakana AI – Fugu Max and Fugu Ultra v2 2026-09-11
Not single models but orchestrators: you make one API call and the system decides which pool of smaller open and specialist models should do the work. Fugu Max scores best overall on six benchmarks at USD 2 in / 6 out per million tokens, which Sakana puts 40–60% below Sonnet 5 and Kimi K3 on output. Notably, Fugu Ultra v2's pool contains no Claude Fable or GPT-6 Astra. Keine einzelnen Modelle, sondern Orchestratoren: Ein API-Aufruf, und das System entscheidet, welcher Pool kleinerer offener und spezialisierter Modelle die Arbeit übernimmt. Fugu Max erzielt bei sechs Benchmarks das beste Gesamtergebnis, zu USD 2 Eingabe / 6 Ausgabe pro Million Tokens – laut Sakana 40–60% unter Sonnet 5 und Kimi K3 bei der Ausgabe. Bemerkenswert: Der Pool von Fugu Ultra v2 enthält weder Claude Fable noch GPT-6 Astra. Non pas des modèles uniques mais des orchestrateurs : un seul appel d'API, et le système décide quel ensemble de modèles plus petits, ouverts ou spécialisés, fera le travail. Fugu Max obtient le meilleur score global sur six bancs d'essai à 2 USD en entrée / 6 USD en sortie par million de tokens, soit selon Sakana 40 à 60% de moins que Sonnet 5 et Kimi K3 en sortie. À noter : l'ensemble de Fugu Ultra v2 ne contient ni Claude Fable ni GPT-6 Astra. Non modelli singoli ma orchestratori: una sola chiamata API e il sistema decide quale insieme di modelli più piccoli, aperti o specializzati, svolgerà il lavoro. Fugu Max ottiene il miglior punteggio complessivo su sei benchmark a 2 USD in ingresso / 6 in uscita per milione di token, secondo Sakana il 40–60% in meno di Sonnet 5 e Kimi K3 in uscita. Da notare: l'insieme di Fugu Ultra v2 non contiene né Claude Fable né GPT-6 Astra. 它们不是单一模型,而是编排器:你只发一次 API 调用,系统自行决定由哪一组更小的开源与专用模型来完成工作。Fugu Max 在六项基准上取得最佳总分,价格为每百万 token 输入 2 美元 / 输出 6 美元,Sakana 称其输出价格比 Sonnet 5 与 Kimi K3 低 40–60%。值得注意的是,Fugu Ultra v2 的模型池中既无 Claude Fable 也无 GPT-6 Astra。
Source →
OpenAI – GPT-Live-1 in the API 2026-09-10
A full-duplex voice model – it listens and speaks at the same time, rather than waiting for you to finish – now callable by developers at USD 0.05 per minute, with 12 new voices, telephony support and built-in transcripts. Makes phone-based data collection technically straightforward; the consent and privacy questions remain the hard part. Ein Vollduplex-Sprachmodell – es hört und spricht gleichzeitig, statt auf das Ende Ihres Satzes zu warten – ist nun für Entwickelnde zu USD 0,05 pro Minute nutzbar, mit 12 neuen Stimmen, Telefonie-Unterstützung und integrierten Transkripten. Damit wird telefonische Datenerhebung technisch einfach; schwierig bleiben Einwilligung und Datenschutz. Un modèle vocal en duplex intégral – il écoute et parle en même temps, au lieu d'attendre la fin de votre phrase – est désormais accessible aux développeurs à 0,05 USD la minute, avec 12 nouvelles voix, la prise en charge de la téléphonie et des transcriptions intégrées. La collecte de données par téléphone devient techniquement simple ; le consentement et la confidentialité restent le point difficile. Un modello vocale full-duplex – ascolta e parla contemporaneamente, invece di attendere la fine della frase – è ora utilizzabile dagli sviluppatori a 0,05 USD al minuto, con 12 nuove voci, supporto telefonico e trascrizioni integrate. Rende tecnicamente semplice la raccolta dati telefonica; restano difficili consenso e privacy. 一款全双工语音模型——可同时聆听与说话,无需等待对方说完——现已开放给开发者,费用为每分钟 0.05 美元,配有 12 种新语音、电话接入支持和内置转写。这让电话数据采集在技术上变得简单;难点仍在于知情同意与隐私。
Source →
DeepSeek – V4.1 Flash 2026-09-10
DeepSeek's smallest new-architecture model, with image understanding built into the model rather than bolted on afterwards. It retires both V4 Flash and V4 Flash Vision Exp, and Artificial Analysis records it as open-weights, so it can in principle be self-hosted on institutional hardware – relevant where data cannot leave the building. DeepSeeks kleinstes Modell der neuen Architektur, mit von Anfang an integriertem Bildverständnis statt nachträglich angebauter Vision. Es ersetzt V4 Flash und V4 Flash Vision Exp; Artificial Analysis führt es als Modell mit offenen Gewichten, es lässt sich also grundsätzlich auf institutionseigener Hardware betreiben – relevant, wenn Daten das Haus nicht verlassen dürfen. Le plus petit modèle de la nouvelle architecture de DeepSeek, avec la compréhension d'images intégrée au modèle plutôt qu'ajoutée après coup. Il remplace V4 Flash et V4 Flash Vision Exp ; Artificial Analysis le recense comme modèle à poids ouverts, il peut donc en principe être hébergé sur du matériel institutionnel – utile lorsque les données ne peuvent pas sortir. Il modello più piccolo della nuova architettura di DeepSeek, con comprensione delle immagini integrata nel modello anziché aggiunta in seguito. Sostituisce V4 Flash e V4 Flash Vision Exp; Artificial Analysis lo registra come modello a pesi aperti, quindi in principio può essere ospitato su hardware dell'istituzione – utile quando i dati non possono uscire. DeepSeek 新架构中体量最小的模型,图像理解为原生内置而非后期加装。它取代了 V4 Flash 与 V4 Flash Vision Exp;Artificial Analysis 将其记录为开放权重模型,原则上可在机构自有硬件上自托管——在数据不得外传的场景下尤为相关。
Source →
Price and plan changes
Anthropic – Claude Code weekly limits reset
A permanent 25% increase over the pre-May baseline took effect on 14 September, replacing a temporary 50% boost that ended on 13 September. Both percentages are accurate against different baselines, but in practice most users now have roughly 17% less weekly headroom than the week before. Worth checking against your team's actual usage before the next billing cycle. Am 14. September trat eine dauerhafte Erhöhung um 25% gegenüber dem Stand vor Mai in Kraft und ersetzte eine temporäre Anhebung um 50%, die am 13. September endete. Beide Prozentzahlen stimmen – nur bezogen auf unterschiedliche Ausgangswerte; praktisch haben die meisten Nutzenden nun rund 17% weniger Wochenbudget als in der Woche davor. Vor dem nächsten Abrechnungszyklus mit der tatsächlichen Team-Nutzung abgleichen. Une augmentation permanente de 25% par rapport au niveau d'avant mai est entrée en vigueur le 14 septembre, remplaçant une hausse temporaire de 50% qui a pris fin le 13 septembre. Les deux pourcentages sont exacts, mais par rapport à des références différentes : en pratique, la plupart des utilisateurs disposent désormais d'environ 17% de marge hebdomadaire en moins que la semaine précédente. À vérifier au regard de l'usage réel de votre équipe avant le prochain cycle de facturation. Il 14 settembre è entrato in vigore un aumento permanente del 25% rispetto al livello precedente a maggio, che sostituisce un incremento temporaneo del 50% concluso il 13 settembre. Entrambe le percentuali sono corrette, ma rispetto a basi diverse: in pratica la maggior parte degli utenti ha ora circa il 17% di margine settimanale in meno rispetto alla settimana precedente. Da verificare con l'uso effettivo del vostro team prima del prossimo ciclo di fatturazione. 相对于五月前基准的永久性 25% 上调已于 9 月 14 日生效,取代了 9 月 13 日结束的临时 50% 提额。两个百分比各自相对不同基准都成立,但实际上多数用户的每周额度比前一周约少 17%。建议在下一个计费周期前,结合团队实际用量核对。
Source →
DeepSeek – V4 Pro stays in service
DeepSeek had said every deepseek-v4-pro request would be answered by V4.1 Flash from 14 September. The API documentation now states that V4 Pro service continues past that date with billing unchanged – so pipelines pinned to that model name do not need rewriting yet. DeepSeek hatte angekündigt, dass ab 14. September jede deepseek-v4-pro-Anfrage von V4.1 Flash beantwortet wird. Die API-Dokumentation besagt nun, dass der V4-Pro-Dienst über dieses Datum hinaus weiterläuft, bei unveränderter Abrechnung – Pipelines, die auf diesen Modellnamen festgelegt sind, müssen also noch nicht umgeschrieben werden. DeepSeek avait annoncé que toute requête deepseek-v4-pro serait traitée par V4.1 Flash à partir du 14 septembre. La documentation de l'API indique désormais que le service V4 Pro se poursuit au-delà de cette date, sans changement de facturation – les chaînes de traitement fixées sur ce nom de modèle n'ont donc pas encore besoin d'être réécrites. DeepSeek aveva annunciato che dal 14 settembre ogni richiesta deepseek-v4-pro sarebbe stata servita da V4.1 Flash. La documentazione dell'API ora indica che il servizio V4 Pro continua oltre tale data, con fatturazione invariata – le pipeline vincolate a quel nome di modello non devono ancora essere riscritte. DeepSeek 曾表示自 9 月 14 日起,所有 deepseek-v4-pro 请求都将由 V4.1 Flash 处理。API 文档现说明 V4 Pro 服务在该日期之后继续提供,计费方式不变——因此锁定该模型名称的流水线暂时无需改写。
Source →
DeepSeek – V4.1 Flash cuts API prices 2026-09-10
Published rates for the new Flash tier are USD 0.15–0.30 per million input tokens on a cache miss, USD 0.003–0.006 on a cache hit, and USD 0.60–1.20 per million output tokens. The lower figure in each pair is the off-peak rate: peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays, so European working hours fall largely inside peak. Die veröffentlichten Sätze für die neue Flash-Stufe liegen bei USD 0,15–0,30 pro Million Eingabe-Tokens bei Cache-Miss, USD 0,003–0,006 bei Cache-Hit und USD 0,60–1,20 pro Million Ausgabe-Tokens. Der jeweils niedrigere Wert ist der Nebenzeiten-Tarif: Hauptzeit ist wochentags 01:00–04:00 und 06:00–10:00 UTC, europäische Arbeitszeiten fallen also überwiegend in die Hauptzeit. Les tarifs publiés pour le nouveau palier Flash sont de 0,15 à 0,30 USD par million de tokens d'entrée en cas d'échec de cache, 0,003 à 0,006 USD en cas de succès de cache, et 0,60 à 1,20 USD par million de tokens de sortie. Le chiffre le plus bas de chaque paire est le tarif heures creuses : les heures pleines sont 01h00–04h00 et 06h00–10h00 UTC en semaine, donc les horaires de travail européens tombent largement en heures pleines. Le tariffe pubblicate per il nuovo livello Flash sono 0,15–0,30 USD per milione di token in ingresso in caso di cache miss, 0,003–0,006 USD in caso di cache hit e 0,60–1,20 USD per milione di token in uscita. Il valore più basso di ogni coppia è la tariffa fuori punta: le ore di punta sono 01:00–04:00 e 06:00–10:00 UTC nei giorni lavorativi, quindi l'orario di lavoro europeo ricade in gran parte nelle ore di punta. 新 Flash 档位的公布价格为:缓存未命中时每百万输入 token 0.15–0.30 美元,缓存命中时 0.003–0.006 美元,每百万输出 token 0.60–1.20 美元。每组中较低者为非高峰价:高峰时段为工作日 UTC 01:00–04:00 与 06:00–10:00,因此欧洲工作时间大多落在高峰段。
Source →
Privacy, rules and terms
Anthropic – Detecting and countering misuse of AI 2026-09-10
Anthropic's threat intelligence team reports disrupting misuse of Claude across seven harm categories between December 2025 and August 2026, including multi-agent intrusion systems running with little human supervision, and stolen API keys resold and used at the victim's expense. The practical lesson for institutions: treat AI API keys as production secrets, and route work only through authorised channels. Anthropics Threat-Intelligence-Team berichtet, zwischen Dezember 2025 und August 2026 Missbrauch von Claude in sieben Schadenskategorien unterbrochen zu haben – darunter Multi-Agenten-Systeme für Einbrüche mit kaum menschlicher Aufsicht sowie gestohlene API-Schlüssel, die weiterverkauft und auf Kosten der Opfer genutzt wurden. Die praktische Lehre für Institutionen: API-Schlüssel wie Produktionsgeheimnisse behandeln und Arbeit nur über autorisierte Kanäle laufen lassen. L'équipe de renseignement sur les menaces d'Anthropic indique avoir déjoué des usages malveillants de Claude dans sept catégories de préjudice entre décembre 2025 et août 2026, dont des systèmes d'intrusion multi-agents fonctionnant avec peu de supervision humaine et des clés d'API volées, revendues et utilisées aux frais de la victime. La leçon pratique pour les institutions : traiter les clés d'API comme des secrets de production et ne faire passer le travail que par des canaux autorisés. Il team di threat intelligence di Anthropic riferisce di aver interrotto usi impropri di Claude in sette categorie di danno tra dicembre 2025 e agosto 2026, tra cui sistemi di intrusione multi-agente operanti con scarsa supervisione umana e chiavi API rubate, rivendute e usate a spese della vittima. La lezione pratica per le istituzioni: trattare le chiavi API come segreti di produzione e far passare il lavoro solo da canali autorizzati. Anthropic 的威胁情报团队报告称,2025 年 12 月至 2026 年 8 月期间,已阻断七类对 Claude 的滥用行为,包括在极少人工监督下运行的多智能体入侵系统,以及被盗后转售、由受害者承担费用的 API 密钥。对机构的实际启示:把 AI API 密钥当作生产环境机密来管理,并且只通过获授权的渠道开展工作。
Source →
California – child-safety chatbot and social-media laws signed 2026-09-10
Governor Newsom signed bipartisan legislation his office describes as the strongest companion-chatbot rules in the United States: new safeguards for companion chatbots, a ban on addictive features such as autoplay and profile-based feeds for under-16s, and wider children's privacy protections. A precedent worth watching if your project puts any conversational tool in front of minors. Gouverneur Newsom hat ein parteiübergreifendes Gesetzespaket unterzeichnet, das sein Büro als strengste Regeln für Begleit-Chatbots in den USA bezeichnet: neue Schutzvorgaben für Companion-Chatbots, Verbot suchterzeugender Funktionen wie Autoplay oder profilbasierter Feeds für unter 16-Jährige sowie erweiterter Kinderdatenschutz. Ein Präzedenzfall, den man beobachten sollte, wenn im Projekt ein Dialogsystem für Minderjährige zum Einsatz kommt. Le gouverneur Newsom a signé une loi bipartisane que son cabinet présente comme l'encadrement des chatbots de compagnie le plus strict des États-Unis : nouvelles obligations de protection pour ces chatbots, interdiction des fonctions addictives comme la lecture automatique ou les fils fondés sur le profil pour les moins de 16 ans, et protection élargie de la vie privée des enfants. Un précédent à suivre si votre projet met un outil conversationnel entre les mains de mineurs. Il governatore Newsom ha firmato una legge bipartisan che il suo ufficio definisce la disciplina più severa degli Stati Uniti per i chatbot di compagnia: nuove tutele per questi chatbot, divieto di funzioni che creano dipendenza come l'autoplay o i feed basati sul profilo per i minori di 16 anni, e protezioni della privacy dei minori ampliate. Un precedente da seguire se il vostro progetto mette uno strumento conversazionale in mano a minori. 纽森州长签署了一项跨党派法案,其办公室称之为全美最严格的陪伴型聊天机器人规则:对陪伴型聊天机器人设定新的保护要求,禁止向 16 岁以下用户提供自动播放、基于画像的信息流等成瘾性功能,并扩大儿童隐私保护。若你的项目会向未成年人提供对话式工具,这是值得关注的先例。
Source →
New features and upgrades
Anthropic – Messages API can compact a conversation on demand
A beta parameter makes the API summarise earlier messages into a signed block that you send in their place on later requests. It addresses the practical ceiling on long agent runs: instead of the conversation growing until it hits the context limit, you compact it deliberately at a point you choose. Ein Beta-Parameter lässt die API frühere Nachrichten zu einem signierten Block zusammenfassen, den man in späteren Anfragen an deren Stelle sendet. Das adressiert die praktische Obergrenze langer Agentenläufe: Statt dass das Gespräch wächst, bis es an das Kontextlimit stösst, verdichtet man es bewusst an einer selbst gewählten Stelle. Un paramètre en version bêta permet à l'API de résumer les messages antérieurs en un bloc signé que vous envoyez à leur place lors des requêtes suivantes. Cela répond au plafond pratique des longues exécutions d'agents : au lieu de laisser la conversation croître jusqu'à la limite de contexte, vous la compactez délibérément au moment choisi. Un parametro in beta consente all'API di riassumere i messaggi precedenti in un blocco firmato da inviare al loro posto nelle richieste successive. Affronta il limite pratico delle lunghe esecuzioni di agenti: invece di lasciare crescere la conversazione fino al limite di contesto, la si compatta deliberatamente nel punto scelto. 一个测试版参数可让 API 把较早的消息汇总为一个带签名的区块,在后续请求中以该区块替代原消息。这解决了长时间智能体运行的实际上限问题:不必让对话一直增长到触及上下文上限,而是在自选的时点主动压缩。
Source →
OpenAI – Agents API in public beta 2026-09-10
The harness that runs Codex is now a hosted product open to all developers: OpenAI manages session orchestration, automatic context compaction and recovery, loads tool definitions only when needed, and offers sandboxes through nine partners. No extra fee – you pay for the tokens and tools your agents use. Das Gerüst, auf dem Codex läuft, ist nun ein gehostetes Produkt für alle Entwickelnden: OpenAI übernimmt Sitzungsorchestrierung, automatische Kontextverdichtung und Wiederaufnahme, lädt Werkzeugdefinitionen nur bei Bedarf und bietet Sandboxes über neun Partner. Keine Zusatzgebühr – bezahlt werden die Tokens und Werkzeuge, die die Agenten nutzen. Le cadre d'exécution de Codex devient un produit hébergé ouvert à tous les développeurs : OpenAI gère l'orchestration des sessions, le compactage automatique du contexte et la reprise, ne charge les définitions d'outils qu'au besoin et propose des bacs à sable via neuf partenaires. Pas de frais supplémentaires – vous payez les tokens et les outils utilisés par vos agents. L'impalcatura su cui gira Codex diventa un prodotto ospitato aperto a tutti gli sviluppatori: OpenAI gestisce l'orchestrazione delle sessioni, la compattazione automatica del contesto e il ripristino, carica le definizioni degli strumenti solo quando serve e offre sandbox tramite nove partner. Nessun costo aggiuntivo – si pagano i token e gli strumenti usati dagli agenti. 支撑 Codex 运行的框架现已成为面向所有开发者的托管产品:OpenAI 负责会话编排、自动上下文压缩与恢复,仅在需要时加载工具定义,并通过九家合作方提供沙箱。不收额外费用——只按智能体所用的 token 和工具计费。
Source →
Anthropic – Managed Agents gain an auto permission policy 2026-09-10
The server can now evaluate each agent or MCP tool call and run it, deny it, or pause for human approval, reporting the decision in the event stream. A new ant beta:sessions connect command attaches your terminal to a running session so you can follow it live and approve or refuse calls as they arrive – useful when an agent touches anything sensitive. Der Server kann nun jeden Agenten- oder MCP-Werkzeugaufruf bewerten und ausführen, ablehnen oder zur menschlichen Freigabe anhalten; die Entscheidung wird im Ereignisstrom mitgeteilt. Ein neuer Befehl ant beta:sessions connect verbindet das Terminal mit einer laufenden Sitzung, sodass man live mitlesen und Aufrufe freigeben oder ablehnen kann – nützlich, wenn ein Agent Sensibles berührt. Le serveur peut désormais évaluer chaque appel d'outil d'agent ou MCP et l'exécuter, le refuser ou le suspendre en attente d'une approbation humaine, la décision étant rapportée dans le flux d'événements. Une nouvelle commande ant beta:sessions connect rattache votre terminal à une session en cours pour la suivre en direct et approuver ou refuser les appels au fil de l'eau – utile lorsqu'un agent touche à des éléments sensibles. Il server può ora valutare ogni chiamata di strumento dell'agente o MCP ed eseguirla, negarla o sospenderla in attesa di approvazione umana, riportando la decisione nel flusso di eventi. Un nuovo comando ant beta:sessions connect collega il terminale a una sessione in corso per seguirla in diretta e approvare o rifiutare le chiamate man mano – utile quando un agente tocca dati sensibili. 服务端现在可对每一次智能体或 MCP 工具调用进行评估,并选择执行、拒绝或暂停以等待人工批准,决策结果会在事件流中报告。新增的 ant beta:sessions connect 命令可将终端接入正在运行的会话,实时跟踪并逐条批准或拒绝调用——当智能体涉及敏感内容时很有用。
Source →
OpenAI – ChatGPT for Financial Services 2026-09-10
A sector-specific ChatGPT built with Morgan Stanley and Evercore, bundling licensed data from Daloopa, PitchBook, LSEG News and Crunchbase for research and client material. The pattern matters more than the sector: expect vertical editions with built-in licensed data sources to appear for other regulated fields. Eine branchenspezifische ChatGPT-Variante, entwickelt mit Morgan Stanley und Evercore, mit lizenzierten Daten von Daloopa, PitchBook, LSEG News und Crunchbase für Recherche und Kundenunterlagen. Wichtiger als die Branche ist das Muster: Mit eingebauten Lizenzdatenquellen ausgestattete Vertikalversionen dürften auch in anderen regulierten Feldern erscheinen. Une version sectorielle de ChatGPT conçue avec Morgan Stanley et Evercore, intégrant des données sous licence de Daloopa, PitchBook, LSEG News et Crunchbase pour la recherche et les documents clients. Le schéma compte plus que le secteur : attendez-vous à des éditions verticales avec sources de données sous licence intégrées dans d'autres domaines réglementés. Una versione settoriale di ChatGPT sviluppata con Morgan Stanley ed Evercore, che integra dati in licenza da Daloopa, PitchBook, LSEG News e Crunchbase per ricerca e materiali per i clienti. Conta più lo schema che il settore: è probabile che compaiano edizioni verticali con fonti dati in licenza integrate anche in altri ambiti regolamentati. 与摩根士丹利和 Evercore 共同打造的行业专版 ChatGPT,内置来自 Daloopa、PitchBook、LSEG News 和 Crunchbase 的授权数据,用于研究与客户材料。比行业本身更值得关注的是这一模式:预计其他受监管领域也会出现内置授权数据源的垂直版本。
Source →
Company news
Microsoft – Suleyman essay on 'model welfare'
Microsoft AI's chief executive published 'A warning about "model welfare"', arguing that AI systems are not conscious and that Anthropic writing uncertainty about Claude's moral status into its constitution is circular reasoning that encourages anthropomorphisation. Anthropic has not published a reply. Worth reading if you advise colleagues on how to talk about what these systems are. Der CEO von Microsoft AI veröffentlichte „A warning about ‚model welfare'" und argumentiert, KI-Systeme seien nicht bewusst und Anthropics Verankerung der Unsicherheit über Claudes moralischen Status in dessen Verfassung sei ein Zirkelschluss, der Vermenschlichung fördere. Anthropic hat bislang nicht geantwortet. Lesenswert, wenn Sie Kolleginnen und Kollegen beraten, wie über diese Systeme zu sprechen ist. Le directeur général de Microsoft AI a publié « A warning about "model welfare" », soutenant que les systèmes d'IA ne sont pas conscients et que l'inscription par Anthropic de l'incertitude sur le statut moral de Claude dans sa constitution relève d'un raisonnement circulaire qui encourage l'anthropomorphisation. Anthropic n'a pas répondu. À lire si vous conseillez des collègues sur la manière de parler de ce que sont ces systèmes. L'amministratore delegato di Microsoft AI ha pubblicato «A warning about "model welfare"», sostenendo che i sistemi di IA non sono coscienti e che l'inserimento da parte di Anthropic dell'incertezza sullo status morale di Claude nella sua costituzione è un ragionamento circolare che incoraggia l'antropomorfizzazione. Anthropic non ha replicato. Da leggere se consigliate i colleghi su come parlare di ciò che questi sistemi sono. 微软 AI 首席执行官发表《A warning about "model welfare"》,主张 AI 系统并无意识,而 Anthropic 把「Claude 道德地位不确定」写入其宪章属于循环论证,会助长拟人化。Anthropic 尚未回应。若你需要就「该如何谈论这些系统的本质」向同事提供建议,值得一读。
Source →
Anthropic – Nvidia reported in talks over IPO stake rumour 2026-09-11
Reuters reports that Nvidia is in talks to commit up to USD 10 billion as anchor investor in an Anthropic listing that could raise as much as USD 100 billion, at a valuation reported at up to USD 2.3 trillion, targeted to price before the US midterm elections in November. Neither company has confirmed the talks. Reuters berichtet, Nvidia verhandle über bis zu USD 10 Milliarden als Ankerinvestor bei einem Börsengang von Anthropic, der bis zu USD 100 Milliarden einbringen könnte, bei einer kolportierten Bewertung von bis zu USD 2,3 Billionen und einem Termin vor den US-Zwischenwahlen im November. Keines der beiden Unternehmen hat die Gespräche bestätigt. Reuters rapporte que Nvidia négocie un engagement pouvant atteindre 10 milliards USD en tant qu'investisseur de référence dans une introduction en bourse d'Anthropic susceptible de lever jusqu'à 100 milliards USD, pour une valorisation évoquée jusqu'à 2 300 milliards USD, prévue avant les élections de mi-mandat américaines de novembre. Aucune des deux entreprises n'a confirmé ces discussions. Reuters riferisce che Nvidia è in trattativa per impegnare fino a 10 miliardi di USD come investitore di riferimento in una quotazione di Anthropic che potrebbe raccogliere fino a 100 miliardi di USD, a una valutazione riportata fino a 2.300 miliardi di USD, prevista prima delle elezioni di metà mandato statunitensi di novembre. Nessuna delle due aziende ha confermato i colloqui. 路透社报道,英伟达正在洽谈以最高 100 亿美元作为基石投资者参与 Anthropic 上市,该次发行或募资至多 1,000 亿美元,估值据报道最高达 2.3 万亿美元,计划在 11 月美国中期选举前定价。两家公司均未确认相关谈判。
Source →
Trending GitHub AI tools
alibaba/open-code-review ★ 32,995 (+3,231)
Alibaba's internal code-review tool, now public: deterministic static checks run alongside an LLM agent, so review comments are grounded in analysis rather than free association. Alibabas internes Code-Review-Werkzeug, nun öffentlich: deterministische statische Prüfungen laufen neben einem LLM-Agenten, sodass Review-Kommentare auf Analyse beruhen statt auf freier Assoziation. L'outil interne de revue de code d'Alibaba, désormais public : des vérifications statiques déterministes s'exécutent aux côtés d'un agent LLM, de sorte que les commentaires reposent sur une analyse plutôt que sur une libre association. Lo strumento interno di revisione del codice di Alibaba, ora pubblico: controlli statici deterministici girano insieme a un agente LLM, così i commenti si basano sull'analisi e non sulla libera associazione. 阿里巴巴内部的代码审查工具现已公开:确定性静态检查与 LLM 智能体并行运行,使审查意见基于分析而非自由联想。
View on GitHub →
JustVugg/colibri ★ 35,272 (+1,546)
Runs large mixture-of-experts models in pure C with no dependencies – aimed at getting frontier-class weights onto hardware you already own rather than rented GPUs. Führt grosse Mixture-of-Experts-Modelle in reinem C ohne Abhängigkeiten aus – mit dem Ziel, Gewichte der Spitzenklasse auf bereits vorhandener Hardware statt auf gemieteten GPUs zu betreiben. Exécute de grands modèles à mélange d'experts en C pur, sans dépendances – pour faire tourner des poids de classe frontière sur du matériel que vous possédez déjà plutôt que sur des GPU loués. Esegue grandi modelli mixture-of-experts in C puro senza dipendenze – per far girare pesi di classe frontiera su hardware che già possedete anziché su GPU noleggiate. 用纯 C、零依赖运行大型专家混合(MoE)模型——目标是让前沿级权重跑在你已有的硬件上,而非租用的 GPU。
View on GitHub →
Tencent/WeKnora ★ 25,752 (+1,197)
A document-knowledge platform pairing retrieval-augmented generation with a reasoning agent – the shape most institutional 'ask our own documents' projects end up needing. Eine Dokumenten-Wissensplattform, die Retrieval-Augmented Generation mit einem Reasoning-Agenten verbindet – genau die Form, die institutionelle Projekte vom Typ „unsere eigenen Dokumente befragen" am Ende brauchen. Une plateforme de connaissance documentaire associant génération augmentée par recherche et agent de raisonnement – la forme dont finissent par avoir besoin la plupart des projets institutionnels de type « interroger nos propres documents ». Una piattaforma di conoscenza documentale che unisce generazione aumentata dal recupero e un agente di ragionamento – la forma di cui finiscono per avere bisogno molti progetti istituzionali del tipo «interrogare i nostri documenti». 一个文档知识平台,将检索增强生成与推理智能体结合——多数机构级「查询我们自己的文档」项目最终需要的正是这种形态。
View on GitHub →
affaan-m/ECC ★ 260,626 (+1,057)
An optimisation harness for coding agents including Claude Code – it tunes how the agent is prompted and orchestrated rather than changing the underlying model. Ein Optimierungsgerüst für Coding-Agenten einschliesslich Claude Code – es justiert Prompting und Orchestrierung des Agenten, ohne das zugrunde liegende Modell zu ändern. Un cadre d'optimisation pour agents de codage, dont Claude Code – il ajuste la manière dont l'agent est sollicité et orchestré, sans modifier le modèle sous-jacent. Un'impalcatura di ottimizzazione per agenti di codifica, incluso Claude Code – regola come l'agente viene sollecitato e orchestrato senza modificare il modello sottostante. 面向编码智能体(含 Claude Code)的优化框架——调整智能体的提示与编排方式,而不改动底层模型。
View on GitHub →
alphaXiv/OpenResearch ★ 4,691 (+1,017)
Turns a coding agent into a research agent: literature search, reading and synthesis driven from the same terminal loop normally used for code. Verwandelt einen Coding-Agenten in einen Forschungs-Agenten: Literatursuche, Lesen und Synthese gesteuert aus derselben Terminal-Schleife, die normalerweise für Code dient. Transforme un agent de codage en agent de recherche : recherche bibliographique, lecture et synthèse pilotées depuis la même boucle de terminal habituellement utilisée pour le code. Trasforma un agente di codifica in un agente di ricerca: ricerca bibliografica, lettura e sintesi guidate dallo stesso ciclo di terminale usato di norma per il codice. 把编码智能体变成研究智能体:文献检索、阅读与综合,均在平时用于写代码的同一终端循环中驱动。
View on GitHub →
cloudflare/security-audit-skill ★ 8,236 (+927)
Cloudflare's multi-phase security-audit skill for coding agents – a vendor shipping its review methodology as an executable agent skill rather than a PDF. Cloudflares mehrphasige Sicherheitsaudit-Fähigkeit für Coding-Agenten – ein Anbieter, der seine Prüfmethodik als ausführbare Agenten-Fähigkeit statt als PDF ausliefert. La compétence d'audit de sécurité multi-phases de Cloudflare pour agents de codage – un fournisseur qui livre sa méthodologie de revue sous forme de compétence exécutable plutôt qu'en PDF. La skill di audit di sicurezza multi-fase di Cloudflare per agenti di codifica – un fornitore che consegna la propria metodologia di revisione come skill eseguibile invece che come PDF. Cloudflare 面向编码智能体的多阶段安全审计技能——厂商以可执行的智能体技能交付其审查方法论,而非一份 PDF。
View on GitHub →