<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>AI Frontier Model Tracker - New Releases</title>
    <description>New frontier AI model releases tracked by DemandSphere. Benchmarks, pricing, and capabilities from OpenAI, Anthropic, Google, xAI, Meta, DeepSeek, and more.</description>
    <link>https://www.demandsphere.com/research/demandsphere-radar/ai-frontier-model-tracker/releases/</link>
    <atom:link href="https://www.demandsphere.com/research/demandsphere-radar/ai-frontier-model-tracker/releases/feed.xml" rel="self" type="application/rss+xml"/>
    <language>en-us</language>
    <lastBuildDate>Thu, 03 Sep 2026 07:06:34 +0000</lastBuildDate>
    <copyright>CC BY-NC 4.0 Ginzamarkets, Inc. (d/b/a DemandSphere)</copyright>
    
    
    
    <item>
      <title>Muse Spark 1.3 (Meta) - Reasoning</title>
      <description>Meta&apos;s strongest model to date and the first that credibly reaches the frontier tier. Artificial Analysis puts Muse Spark 1.3 (max) at 62 on the Intelligence Index, third overall behind Claude Fable 5.1 and Claude Opus 5, and trading blows with GPT-5.6 Sol - independent scoring, not vendor-run. The efficiency claim is the interesting part: 20% fewer tool calls and 25% fewer tokens consumed than Muse Spark 1.2 for the same work. Built for long-horizon sessions, juggling multiple workflows in one thread, generating its own context across conflicting sources, and correcting gaps in its own plan. Priced the same as 1.2 at $1.25/$4.25 per 1M. The 62 figure is for the max tier, which is in limited preview for Meta partners rather than generally available. Open weights promised, no date. Context: 1000K tokens. Pricing: $1.25/M input, $4.25/M output.</description>
      <link>https://research.meta.ai/blog/introducing-muse-spark-1-3</link>
      <guid isPermaLink="false">ds-tracker-ms13-2026-09-02</guid>
      <pubDate>Wed, 02 Sep 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>Meta</category>
    </item>
    
    
    <item>
      <title>Gemini 3.8 Flash Cyber (Google) - Reasoning</title>
      <description>The restricted-access sibling of Gemini 3.8 Flash. Google describes the pair as one core model behind two access envelopes, so capability tracks 3.8 Flash while availability is gated for security work. Listed separately on the same basis as Claude Mythos: this tracker records access envelopes as distinct entries. Note that GPT-5.6-Cyber was excluded in August under the older reading that Cyber builds are specialised variants; that entry is worth revisiting for consistency. Context: 1000K tokens. Pricing: $0.75/M input, $3.75/M output.</description>
      <link>https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash</link>
      <guid isPermaLink="false">ds-tracker-g38fc-2026-09-02</guid>
      <pubDate>Wed, 02 Sep 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>Google</category>
    </item>
    
    
    <item>
      <title>Gemini 3.8 Flash (Google) - Reasoning</title>
      <description>Google&apos;s fourth Flash release in under four months. Terminal-Bench 2.1 90.8%, up from 3.7 Flash&apos;s 81.6%, and it beats Claude Opus 5 on three of the benchmarks Google published. Artificial Analysis Intelligence Index 59 at high reasoning, up 3 from 3.7 Flash. 1M-token context, 64K output, text/image/audio/video/PDF in. Two caveats from Google&apos;s own documentation: 3.8 Flash is built on 3.7 Flash rather than a new base model, and it deliberately spends more thinking tokens - Google explicitly recommends staying on 3.7 Flash for efficiency-first workloads. Pricing $0.75/$3.75 per 1M through December 31, 2026, doubling to $1.50/$7.50 on January 1, so the launch rate is promotional. Batch and Flex are half those rates, Priority is 1.8x. Context: 1000K tokens. Pricing: $0.75/M input, $3.75/M output.</description>
      <link>https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash</link>
      <guid isPermaLink="false">ds-tracker-g38f-2026-09-02</guid>
      <pubDate>Wed, 02 Sep 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>Google</category>
    </item>
    
    
    <item>
      <title>Claude Mythos 5.1 (Anthropic) - Reasoning</title>
      <description>The same underlying model as Fable 5.1 at a different safeguard level: Fable 5.1 is generally available, Mythos 5.1 sits behind trusted-access programs for cybersecurity and life-sciences work. Benchmarks and pricing track Fable 5.1. Listed separately because the tracker records access envelopes as distinct entries, the same treatment Mythos 5 received in June. Context: 1000K tokens. Pricing: $10/M input, $50/M output.</description>
      <link>https://www.anthropic.com/news/claude-fable-5-mythos-5</link>
      <guid isPermaLink="false">ds-tracker-cm51-2026-09-01</guid>
      <pubDate>Tue, 01 Sep 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>Anthropic</category>
    </item>
    
    
    <item>
      <title>Claude Fable 5.1 (Anthropic) - Reasoning</title>
      <description>Three months after Fable 5, and the jump is in agentic work rather than raw knowledge. Anthropic reports Terminal-Bench-Science 52.6% against Fable 5&apos;s 24.7%, Opus 5&apos;s 29.0%, and GPT-5.6 Sol&apos;s 22.4% - more than double its own predecessor. GDPval-AA v2 1853, ahead of Opus 5 at 1824 and Fable 5 at 1723. OSWorld 2.0 partial pass 77.9% (Opus 5 75.4%, Fable 5 72.9%); strict pass 41.7% against 39.6% and 36.1%. It finishes ahead of Opus 5 on every category Anthropic published. Sticker price unchanged at $10/$50 per 1M; the commercial change is cache reads cut 75%, from $1.00 to $0.25 per 1M, which matters more than the headline for agent workloads that re-read context constantly. Standard benchmark columns left blank - Anthropic published none of them for this release. Context: 1000K tokens. Pricing: $10/M input, $50/M output.</description>
      <link>https://www.anthropic.com/news/claude-fable-5-mythos-5</link>
      <guid isPermaLink="false">ds-tracker-cf51-2026-09-01</guid>
      <pubDate>Tue, 01 Sep 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>Anthropic</category>
    </item>
    
    
    <item>
      <title>Qwen3.8-Flash-Next (Alibaba/Qwen) - Reasoning</title>
      <description>A preview of the Qwen4 architecture rather than a product release - Alibaba shipped it early so tooling, quantization, and pipelines can adapt before the full Qwen4 family lands. 125B MoE with only 6B active per token across 512 experts (10 routed plus one shared), alongside a separate 51B n-gram embedding table and a 4B multi-token prediction layer, for 176B total on disk. 262,144-token native context, extensible to 1M. Multimodal in. Open weights on ModelScope in standard and FP8. Vendor-reported: GPQA Diamond 91.7, SWE-bench Pro 62.5, DeepSWE v1.1 58.7, CoWorkBench 73.9 against DeepSeek-V4-Flash&apos;s 45.1, JobBench 55.7 against Qwen3.7-Plus&apos;s 27.6. The active-parameter count is low enough to run locally in roughly 75GB of RAM, which is the point of the release. Listed here for the architecture rather than the tier - the GPQA figure is frontier-level and the Qwen4 preview status makes it the most informative Qwen release since 3.8-Max. Pricing not published at time of entry. Context: 262K tokens. Open weights.</description>
      <link>https://www.marktechpost.com/2026/08/26/alibabas-qwen-team-releases-qwen3-8-flash-next-a-125b-multimodal-moe-with-6b-active-parameters-previewing-the-qwen4-architecture/</link>
      <guid isPermaLink="false">ds-tracker-qw38fn-2026-08-26</guid>
      <pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Open</category>
      <category>Alibaba/Qwen</category>
    </item>
    
    
    <item>
      <title>GLM-5.3-Flash (Zhipu AI) - Reasoning</title>
      <description>The first natively multimodal model in the GLM-5 family, and a different animal from GLM-5.3 despite the name - 320B MoE with 18B active against GLM-5.3&apos;s 743B/40B, trained on a 30T-token multimodal corpus rather than distilled from the larger model. Linear attention for local dependencies combined with sparse attention for global context. 1M-token window, image and video in. MIT weights on Hugging Face at release, which GLM-5.3 itself still did not have on this date. Z.ai reports Terminal-Bench 2.1 84.3, within a point of Claude Opus 4.8 at 85.0; DeepSWE v1.1 63.4 against GLM-5.2&apos;s 46.2; AutomationBench 48.8 against 26.2; OfficeQA Pro 62.4, ahead of both Opus 4.8 and DeepSeek-V4-Vision-Exp. Artificial Analysis Intelligence Index v4.1.1 puts it at 57. $0.15/$0.50 per 1M with cached input at $0.03, halved through September 9 as a launch promotion. This is the model that ran anonymously as ox-alpha on OpenRouter and OpenCode from August 20; Z.ai confirmed the connection at launch. Context: 1000K tokens. Pricing: $0.15/M input, $0.5/M output. Open weights.</description>
      <link>https://huggingface.co/zai-org/GLM-5.3-Flash</link>
      <guid isPermaLink="false">ds-tracker-glm53f-2026-08-26</guid>
      <pubDate>Wed, 26 Aug 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Open</category>
      <category>Zhipu AI</category>
    </item>
    
    
    <item>
      <title>GLM-5.3 (Zhipu AI) - Reasoning</title>
      <description>Same 743B MoE / roughly 40B active base as GLM-5.2 - every gain here comes from extended post-training rather than a new pre-train, which is the notable part. 1M-token context, up to 128K output. Z.ai reports a 50% coding improvement over GLM-5.2 on its in-house Code Bench (34.5% at Max effort, ~75K output tokens per task; 31.4% at High), Terminal-Bench 3.0 28.3 (from 4.6), DeepSWE v1.1 66.9 (from 46.2), and Agents&apos; Last Exam 28.5 (from 23.8), claiming open-source SOTA on the first two. The bigger jump is in cybersecurity: CyberGym 84.5% (from 77.2%) and ExploitBench 54.4% (from 24.4%), with Z.ai&apos;s own testing surfacing 2,436 vulnerabilities including 1,097 medium- and high-severity. Weights landed on zai-org on August 28 after the safety review: 141 safetensors shards, 756GB, which at that size implies FP8 rather than BF16. Not MIT, unlike GLM-5.2 and GLM-5.3-Flash - the custom GLM-5.3 License permits commercial use, fine-tuning and redistribution, but requires any Model-as-a-Service operator with over $10B in revenue across any 12 consecutive months to pass a Z.ai security review first. That threshold reads as aimed at hyperscalers rather than at ordinary users. Launch was widely reported on August 14; z.ai&apos;s own release notes log it under August 18. Z.ai puts API pricing at roughly a tenth of US frontier rates but has not published per-token figures. Context: 1000K tokens. Open weights.</description>
      <link>https://docs.z.ai/guides/llm/glm-5.3</link>
      <guid isPermaLink="false">ds-tracker-glm53-2026-08-14</guid>
      <pubDate>Fri, 14 Aug 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Open</category>
      <category>Zhipu AI</category>
    </item>
    
    
    <item>
      <title>Gemini 3.7 Flash (Google) - Reasoning</title>
      <description>Google&apos;s Flash refresh just 23 days after Gemini 3.6 Flash, pitched at complex coding, agentic workflows, and reliable multi-step execution. 1,048,576-token input window and up to 65,536 output tokens. Multimodal input (text, image, video, audio, PDF), text output. Introductory pricing of $0.75/$3.75 per 1M tokens runs through December 31, 2026; on January 1, 2027 it doubles to $1.50/$7.50, which is exactly what 3.6 Flash cost - so the launch rate is a promotion, not a price cut. Google reports DeepSWE v1.1 65.3 (from 49.0), FrontierCode 1.1 Main 43.6 (from 34.4), and AutomationBench 30.4 (from 17.0). Artificial Analysis scores it 56 on the Intelligence Index against 52 for 3.6 Flash, and ranks it first of 186 models on output speed at 340.1 tokens per second. Standard benchmark columns left blank - Google did not publish MMLU-Pro, GPQA Diamond, or SWE-bench Verified for this release. Context: 1000K tokens. Pricing: $0.75/M input, $3.75/M output.</description>
      <link>https://ai.google.dev/gemini-api/docs/models/gemini-3.7-flash</link>
      <guid isPermaLink="false">ds-tracker-g37f-2026-08-13</guid>
      <pubDate>Thu, 13 Aug 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>Google</category>
    </item>
    
    
    <item>
      <title>Grok 4.6 (xAI) - Reasoning</title>
      <description>xAI&apos;s flagship refresh barely a month after Grok 4.5, aimed at long-running agents, agentic coding, and visual or interactive work. Artificial Analysis Intelligence Index 61, level with GPT-5.6 Sol and one point behind Claude Fable 5. The gains over 4.5 are concentrated in agentic coding: DeepSWE v1.1 65.9 (from 54.0) and APEX-Agents 57.5 (from 47.1). Other results xAI published: CursorBench v3.2 69.9%, FrontierCode v1.1 61.3%, APEX-SWE 56.4%, Terminal-Bench v3.0 26%, GDPVal-AA v2 1753, AA-Briefcase 1577, Harvey LAB (Vals) 15.8%. On xAI&apos;s own ten-row eval table Claude Fable 5 Max still takes the most first-place rows; Grok 4.6 wins knowledge work and legal. 500K context, knowledge cutoff February 1, 2026. $2/$6 per 1M tokens, with a fast variant at twice the price. Standard benchmark columns left blank - xAI did not publish MMLU-Pro, GPQA Diamond, AIME, or SWE-bench Verified for this release. Context: 500K tokens. Pricing: $2/M input, $6/M output.</description>
      <link>https://x.ai/news/grok-4-6</link>
      <guid isPermaLink="false">ds-tracker-grok46-2026-08-12</guid>
      <pubDate>Wed, 12 Aug 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>xAI</category>
    </item>
    
    
    <item>
      <title>Nemotron 3.5 Lightning (NVIDIA) - Reasoning</title>
      <description>NVIDIA&apos;s first entry in the tracker and the sharpest statement yet of speed over peak intelligence. 30B Mixture-of-Experts with only 3B active per token, on a Mamba2-Transformer hybrid architecture with Multi-Token Prediction - the low active-parameter count is the whole point, buying up to 4x the output speed of similar-sized models. Open weights under OpenMDW-1.1, BF16 and NVFP4 formats on Hugging Face. Context up to 1M tokens, though NVIDIA runs 256K for single-H100 deployment. Pre-trained on 20T+ tokens spanning English, 19 other spoken languages, and 43 programming languages; pre-training cutoff September 2025, post-training May 2026. Reasoning is configurable on or off via the chat template. Text in, text out. Model-card benchmarks: MMLU-Pro 81.94, GPQA Diamond 75.44, SWE-bench Verified 51.56, IFBench (loose) 71.88, AA-LCR 52.00, Terminal-Bench 2.1 24.58. NVIDIA also reports 86% on its own PinchBench while completing 10,000 tasks 30% faster than Qwen3.6 35B. Positioned for high-volume, low-latency agent execution rather than frontier reasoning - read the Terminal-Bench score in that light. Context: 1000K tokens. Open weights.</description>
      <link>https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16</link>
      <guid isPermaLink="false">ds-tracker-nem35l-2026-08-11</guid>
      <pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Open</category>
      <category>NVIDIA</category>
    </item>
    
    
    <item>
      <title>Muse Glimmer (Meta) - General</title>
      <description>Roughly 30B dense causal transformer including the vision tower, released August 10 under Apache 2.0 and built to run on one consumer GPU - quantized to under 20GB so the model and its supporting components fit a 24GB or 32GB memory envelope. Text and image in, text out; audio is unsupported and video is processed as individual frames. Trained on data from more than 100 languages. 131,072-token context, knowledge cutoff January 4, 2026. On Hugging Face with Ollama, LM Studio, llama.cpp, ExecuTorch, and MLX integrations. Meta-published results: AIME 2026 94.7, MCP Atlas 75.5, SWE-Bench Pro 51.2, DeepSearch QA 74.6, AA-LCR 80.0, IFBench 77.0, OSWorld-Verified 65.9, Gaia2 43.3. Meta&apos;s comparison against Gemma 4 31B and Qwen3.6-27B takes the better of each competitor&apos;s self-reported or Meta-reproduced score, so read the head-to-heads accordingly. Positioned as an agentic model rather than a reasoning model. Zuckerberg framed the release around distributing rather than centralizing capability. Context: 131K tokens. Open weights.</description>
      <link>https://www.marktechpost.com/2026/08/10/meta-ai-releases-muse-glimmer/</link>
      <guid isPermaLink="false">ds-tracker-muse-glimmer-2026-08-10</guid>
      <pubDate>Mon, 10 Aug 2026 00:00:00 +0000</pubDate>
      <category>General</category>
      <category>Open</category>
      <category>Meta</category>
    </item>
    
    
    <item>
      <title>Muse Spark 1.2 (Meta) - Reasoning</title>
      <description>Meta&apos;s flagship coding model, released August 5 through the Meta Model API and the model behind Muse Code, a terminal coding agent for macOS and Linux. 1,048,576-token context, unchanged from 1.1. Reasoning is mandatory with five effort levels (minimal, low, medium, high, xhigh; medium default) - there is no non-thinking configuration. Standard tier $1.25/$4.25 per 1M with cached input at $0.15; a contributor tier drops to $0.10/$0.20 in exchange for opting into Meta training, capped at 60 requests/min. Listed for text, image, video, audio, and PDF input, though independent evaluation so far covers only text and image. All published numbers are vendor-run inside Meta&apos;s own harness: Terminal-Bench 2.1 82.9%, DeepSWE v1.1 59.3%, Artificial Analysis Intelligence Index 54 (tied with Grok 4.5). No third party has published scores from a neutral scaffold. Parameter count and knowledge cutoff undisclosed. Meta says weights are coming, without a date or license. Context: 1000K tokens. Pricing: $1.25/M input, $4.25/M output.</description>
      <link>https://www.orcarouter.ai/blog/meta-muse-spark-1-2-explained</link>
      <guid isPermaLink="false">ds-tracker-muse-spark-12-2026-08-05</guid>
      <pubDate>Wed, 05 Aug 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>Meta</category>
    </item>
    
    
    <item>
      <title>Qwen3.8-Max (Alibaba/Qwen) - Reasoning</title>
      <description>Alibaba&apos;s flagship Qwen3.8-Max, previewed July 19 and made generally available via API (QwenCloud) on August 3, with open weights due the following week - Alibaba&apos;s first open-weight model at Max scale, alongside a smaller Qwen3.8-27B open checkpoint. Sparse Mixture-of-Experts, 2.4T total with roughly 95B active per token, 1M context, and Qwen&apos;s first multimodal model above 1T (text, image, and video in; text out). Alibaba positions it against GPT-5.6 Sol and Claude Fable 5. Vendor-reported benchmarks (not independently verified): PaperBench 93.0 vs Fable 5&apos;s 88.8, IFBench 82.8 vs 63.5, and Terminal-Bench 2.1 86.6 - ahead of Opus 4.8 and Fable 5 at 84.6, behind GPT-5.6 Sol at 88.8. Emphasizes coding and long-horizon autonomous operation (10+ days of continuous software development). Per-token API pricing not broken out at launch; the predecessor Qwen3.7-Max ran $2.50/$7.50 per 1M. Context: 1000K tokens.</description>
      <link>https://www.marktechpost.com/2026/08/03/alibaba-qwen-releases-qwen3-8-max/</link>
      <guid isPermaLink="false">ds-tracker-qw38-max-2026-08-03</guid>
      <pubDate>Mon, 03 Aug 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>Alibaba/Qwen</category>
    </item>
    
    
    <item>
      <title>DeepSeek V4 Flash (DeepSeek) - Reasoning</title>
      <description>284B MoE / 13B active. MIT license. Same hybrid attention as V4-Pro. 1M context. $0.14/$0.28 per 1M tokens. Shipped as a preview on April 24 and reached general availability July 31 as DeepSeek-V4-Flash-0731 - not a new architecture but the April checkpoint retrained through a reworked post-training pipeline aimed at coding, agents, reasoning, and tool use. DeepSeek reports the 0731 build beats the much larger V4-Pro preview on every agentic benchmark it published. Weights are MIT-licensed on Hugging Face and supersede the preview checkpoint; the official V4-Flash API entered public beta the same day. Benchmark columns below are from the April preview: 88.1% GPQA Diamond, 91.6% LiveCodeBench, 86.2% MMLU-Pro. Context: 1000K tokens. Pricing: $0.14/M input, $0.28/M output. Open weights.</description>
      <link>https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731</link>
      <guid isPermaLink="false">ds-tracker-dsv4f-2026-07-31</guid>
      <pubDate>Fri, 31 Jul 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Open</category>
      <category>DeepSeek</category>
    </item>
    
    
    <item>
      <title>Claude Opus 5 (Anthropic) - Reasoning</title>
      <description>Anthropic&apos;s new flagship, released July 24 and positioned as a daily-driver Opus that rivals Fable 5 at half the price ($5/$25 vs $10/$50). Anthropic reports it more than doubles Opus 4.8 on Frontier-Bench v0.1 and surpasses all other models there; lands within 0.5% of Fable 5&apos;s peak on CursorBench 3.2 at half the cost; scores roughly 3x the next-best model on ARC-AGI 3; and beats Fable 5&apos;s best OSWorld 2.0 result at about one-third the cost. Standard SWE-bench / GPQA figures were not published at launch. 1M context, multimodal, adaptive thinking. Closed, via the Anthropic API. Context: 1000K tokens. Pricing: $5/M input, $25/M output.</description>
      <link>https://www.anthropic.com/news/claude-opus-5</link>
      <guid isPermaLink="false">ds-tracker-co5-2026-07-24</guid>
      <pubDate>Fri, 24 Jul 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>Anthropic</category>
    </item>
    
    
    <item>
      <title>Gemini 3.6 Flash (Google) - Reasoning</title>
      <description>Google&apos;s mid-tier Gemini refresh, launched July 21 with day-one availability in AI Studio, the Gemini API, and the Gemini app. 1,048,576-token input window and up to 65,536 output tokens. $1.50/$7.50 per 1M tokens with cached input at $0.15 - a lower output rate than Gemini 3.5 Flash despite the capability bump. Multimodal input (text, image, video, audio, PDF), text output. Google reports it uses roughly 17% fewer output tokens than 3.5 Flash while beating it on every published benchmark, including SWE-Bench Pro 58.7 vs 55.1 and OSWorld-Verified 83.0 vs 78.4. Artificial Analysis puts it at 52 on the Intelligence Index at high reasoning effort. Standard benchmark columns left blank - Google did not publish MMLU-Pro, GPQA Diamond, or SWE-bench Verified for this release. Context: 1000K tokens. Pricing: $1.5/M input, $7.5/M output.</description>
      <link>https://openrouter.ai/google/gemini-3.6-flash</link>
      <guid isPermaLink="false">ds-tracker-g36f-2026-07-21</guid>
      <pubDate>Tue, 21 Jul 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>Google</category>
    </item>
    
    
    <item>
      <title>Kimi K3 (Moonshot AI) - Reasoning</title>
      <description>Moonshot AI&apos;s 2.8T-parameter open-weight multimodal reasoning MoE - 896 experts with 16 activated per token, an unusually high sparsity ratio, so active compute sits far below the headline count. Built on Kimi Delta Attention (KDA) hybrid linear attention plus Attention Residuals for efficiency (up to ~6.3x faster decoding at 1M-token context). Native vision, 1M context. Released July 16 via the Kimi apps and API, with full open weights scheduled for July 27. Independent testing places it #4 among all frontier models, trailing only Claude Fable 5 and GPT-5.6 Sol and edging past Opus 4.8; debuted #1 on LMArena&apos;s Frontend Code Arena (1679 Elo) and 1486 on the main text board. Artificial Analysis: Intelligence Index 57.1, Coding 76.2, Agentic 50.1. Pricing $3/$15 per 1M tokens (cache-hit input $0.30). Reasoning effort currently supports only the &apos;max&apos; level. Context: 1000K tokens. Pricing: $3/M input, $15/M output. Open weights.</description>
      <link>https://openrouter.ai/moonshotai/kimi-k3</link>
      <guid isPermaLink="false">ds-tracker-kk3-2026-07-16</guid>
      <pubDate>Thu, 16 Jul 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Open</category>
      <category>Moonshot AI</category>
    </item>
    
    
    <item>
      <title>Inkling (Thinking Machines Lab) - Reasoning</title>
      <description>Thinking Machines Lab&apos;s debut model and first public release (Mira Murati&apos;s team). Sparse Mixture-of-Experts with 975B total / 41B active parameters (256 routed plus 2 shared experts, 6 routed active per token), pretrained on 45T tokens of text, image, audio, and video. 1M context with efficient thinking-mode reasoning. Open weights on Hugging Face, fine-tunable via the lab&apos;s Tinker platform. The lab is candid that it is &apos;not the strongest overall model available today, open or closed,&apos; emphasizing multimodal breadth and efficient reasoning instead. Self-reported benchmarks at effort=0.99 (not independently verified): GPQA Diamond 87.2, AIME 2026 97.1, SWE-bench Verified 77.6, Global-MMLU-Lite 88.7, and HLE with tools 46.0. Served on Tinker (limited-time 50% off, with 64K and 256K context tiers) and via TogetherAI, Fireworks, Modal, Databricks, and Baseten; no first-party per-token pricing disclosed yet. A companion Inkling-Small preview (276B total / 12B active) was announced alongside it. Context: 1000K tokens. Open weights.</description>
      <link>https://thinkingmachines.ai/news/introducing-inkling/</link>
      <guid isPermaLink="false">ds-tracker-inkling-2026-07-15</guid>
      <pubDate>Wed, 15 Jul 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Open</category>
      <category>Thinking Machines Lab</category>
    </item>
    
    
    <item>
      <title>GPT-5.6 Luna (OpenAI) - Reasoning</title>
      <description>Fastest and cheapest tier of the GPT-5.6 family. $1/$6 per 1M tokens. 1M context. Context: 1000K tokens. Pricing: $1/M input, $6/M output.</description>
      <link>https://openai.com/index/previewing-gpt-5-6-sol/</link>
      <guid isPermaLink="false">ds-tracker-gpt-56luna-2026-07-09</guid>
      <pubDate>Thu, 09 Jul 2026 00:00:00 +0000</pubDate>
      <category>Reasoning</category>
      <category>Closed</category>
      <category>OpenAI</category>
    </item>
    
  </channel>
</rss>
