Agent analysis
Standing syntheses the agent writes over each lab's captured pages, structured signals, and bounded web evidence — every material claim cited back to its source.
InclusionAI (Ant Group)InclusionAI (Ant Group)22hInclusionAI (Ant Group) is running a high-cadence, open-weight “efficiency + agentic” strategy: a dense stream of Mixture-of-Experts autoregressive models (Ling), diffusion language models (LLaDA), and reasoning models (Ring), wrapped in vertical agentic finetunes, safety guardrails, and the serving/training/sandbox tooling to run them. Efficiency is the recurring motif—sparse MoE with tiny active-parameter counts…IBM (Granite)IBM (Granite)22hIBM's Granite organization is running two tracks at once: an open-weights enterprise-AI line anchored by the Apache-2.0 Granite 4.2 reasoning family, and a parallel hardware/quantum bet that the model releases feed into a broader “enterprise AI + quantum” story. The shipping pattern is the tell — the same base models appear as Hugging Face weights, Apple-silicon MLX conversions, and FP8/NVFP4/MXFP4 quantized…Cloudflare (Workers AI)Cloudflare (Workers AI)22hCloudflare is running a two-sided bet that the "agentic Internet" is its next platform cycle. On one side, it is selling itself as the connectivity and security fabric for AI agents — governing MCP server sprawl and unmanaged agents via Cloudflare One, reframing the web's business model around AI crawlers and content markets, and pushing isolates/Durable Objects as the compute primitive for agents via…BasetenBaseten22hBaseten is an inference-first neocloud that raised a $1.5B Series F led by Altimeter Capital, Conviction Partners, and Spark Capital. The evidence here shows it now scaling on three fronts at once: enterprise platformization (identity/authorization, read-only roles, AWS IAM AssumeRole); a climb up the stack from inference into post-training/RL infrastructure; and a new research arm, Base Labs, with a…AnthropicAnthropic23hAnthropic's September 2026 signals show a lab running three plays at once: pushing frontier capability into math and science; hardening and publicly disclosing its safety/alignment machinery; and industrializing compute, finance, and go-to-market ahead of a reported IPO. The week's shipping is developer-tooling cadence (Claude Code, agent SDK, Foundation Models) rather than a new flagship model, while the headline…Inception LabsInception Labs1dInception Labs is building a new model category — diffusion LLMs (dLLMs) — and betting that raw generation speed, not peak reasoning depth, is the binding constraint for production AI. Its diffusion architecture generates many tokens in parallel via coarse-to-fine denoising rather than sequential autoregression, which it claims delivers 1,000+ tokens/sec on commodity NVIDIA GPUs (H100/Blackwell) and up to 10x faster…Amazon (Nova)Amazon (Nova)1dAmazon's public research surface reads as a lab in mid-consolidation. The open-science surface keeps shipping at high cadence — time-series forecasting (Chronos-2), agentic evaluation benchmarks, and a rapidly iterating Python concurrency library (concurry) — while the commercial Nova flagship family (Premier, Omni, Reel, Canvas) is being wound down in favor of a Frontier Model Research group led by Pieter Abbeel,…CerebrasCerebras6dCerebras is consolidating into a speed layer for frontier inference, not a model house. The through-line across the pack is that wafer-scale hardware removes the GPU memory wall and the multi-chip networking tax, letting Cerebras serve the same frontier model — GPT-5.6 Sol with the same architecture, weights, precision, context configuration, and reasoning settings — at up to 750 output tokens/second, where "the…DeepSeekDeepSeek1wDeepSeek's center of gravity is shifting from flagship frontier models toward agentic productization and efficiency infrastructure. The dominant recent artifact is DeepSeek Harness ('dsh'), an MIT-licensed TypeScript agent harness built on an 'everything is a plugin' Cordis architecture, which shipped seven tagged prereleases in roughly two weeks (v0.1.0-rc.7 → v0.1.2-alpha.2). Hiring and public framing confirm the…Arcee AIArcee AI1wArcee AI is pivoting from an adapter/fine-tune shop into a vertically integrated, US-based open-weight lab. The evidence shows it now trains frontier-scale sparse Mixture-of-Experts models from scratch (Trinity Nano/Mini/Large), stands up its own training and RL infrastructure (Slurm, NeMo-RL, prime-rl, pybubble), and is productizing both an agent runtime (nac) and a multi-model inference API. The recurring…CohereCohere1wCohere is running a security-first, enterprise-and-sovereign play rather than a consumer-scale model race. The dominant signals in this pack are commercialization and deployment buildout: a dense wave of forward-deployed engineering (FDE) hiring across Infrastructure, Agentic Platform, and Sovereign AI teams, alongside Solutions Architecture, Account Executive, and Customer Success expansion into Nordics, UAE/META,…ClarifaiClarifai1wClarifai's observable footprint is that of a platform being folded into a neocloud, not a frontier model lab: an active multi-language SDK/API release program (Python, Node.js, and gRPC clients across five languages) with essentially no new hiring, no fork or research activity, and no model-card releases in this pack. The decisive signal is ownership: an engineer describes joining "Nebius as part of the Clarifai…Baidu (ERNIE)Baidu (ERNIE)2wBaidu is running two visible plays in this pack. First, a heavy open-source, deployment-and-interoperability push across the PaddlePaddle ecosystem — PaddlePaddle v3.0, PaddleOCR v3.x, PaddleX v3.x, and Paddle2ONNX v2.x — aimed at multi-hardware inference, serving, and model export. Second, a frontier ERNIE effort with open-weights releases (ERNIE-4.5-Thinking, ERNIE-Image, Unlimited-OCR) and sustained…AI21 LabsAI21 Labs3wAI21 Labs is repositioning from a pure foundation-model vendor toward an agent-orchestration and inference-optimization platform. The clearest signal is operational: the company cut headcount from roughly 180 to about 70 people and will discontinue selling standalone language models, refocusing on Maestro, its AI agent management platform. The same pivot shows up in public research and product writing, which now…ByteDance (Doubao/Seed)ByteDance (Doubao/Seed)3wByteDance's Seed team is a full-stack frontier lab, not a single-model shop: it spans foundation LLMs, speech, vision, world models, robotics, agents, and AI infrastructure, with labs across China, Singapore, and the U.S.. It is large and commercially backed — roughly 2,000 employees led by former Google DeepMind research VP Wu Yonghui — and it pairs unusually deep open-source systems releases (VeOmni,…OpenBMB (MiniCPM)OpenBMB (MiniCPM)4wOpenBMB is executing a multi-vector strategy that spans on-device foundation models, embodied AI (VLA robotics), enterprise agent platforms (StaffDeck/PilotDeck), data refinement tooling (UltraX), and HPC optimization (ForgeStencil). The lab operates as a joint academic-industrial consortium — Tsinghua THUNLP, ModelBest, NEU-ModelBest Data Intelligence Joint Lab, and AI9Stars — and systematically open-sources models…LG AI Research (EXAONE)LG AI Research (EXAONE)4wLG AI Research is executing a dual-track strategy: advancing frontier-scale foundation models under Korea's sovereign AI program while simultaneously building a portfolio of industry-specialized Expert AI models for B2B commercialization. The July 2026 release of K-EXAONE 2.0 — a 750B-parameter Mixture-of-Experts model with 37B active parameters, developed under the Ministry of Science and ICT's Sovereign AI…Databricks (DBRX)Databricks (DBRX)4wDatabricks in this evidence window is executing a three-phase enterprise platform consolidation: (1) an internal SAP S/4HANA transformation with an agentic AI layer that doubles as both dogfooding and product blueprint; (2) a platform-infrastructure hardening wave spanning ingestion SDKs, networking, access management, and multimodal data types; and (3) an aggressive GTM scaling motion into European regulated…CoreWeaveCoreWeave4wCoreWeave is executing a deliberate pivot from a pure-play GPU neocloud into a vertically integrated AI platform company. The evidence shows the firm layering a proprietary agent-development platform (W&B Weave), an autonomous research agent (ARIA), a serverless inference service, and agentic sandboxing atop its raw compute footprint—while continuing to scale its infrastructure base aggressively. The CEO's 2025…MiniMaxMiniMax4wMiniMax has evolved from a linear-attention LLM shop into a full-stack, omni-modal AI lab shipping models across text, vision, video, audio, and agentic workflows on a quarterly-or-faster cadence. The lab's public posture is aggressively open-source — releasing weights, papers, and tooling under Apache-2.0 and MIT licenses — while simultaneously building a developer ecosystem through MCP servers, a Vercel AI SDK…MicrosoftMicrosoft4wMicrosoft is signaling a full-stack agentic-AI buildout that spans training environments, agent memory, skill optimization, self-improving inference, and open-weight model releases. The evidence cluster — rapid agent-learning releases, research on computer-use training worlds (Echoverse), evolving knowledge systems (EvoLib), trainable agent skills (SkillOpt), and memory architectures (Memora) — points to a lab…Nous ResearchNous Research4wNous Research is executing a decisive pivot from open-weight model publisher to integrated agent platform company. Throughout mid-2026, the overwhelming majority of engineering velocity has concentrated on hermes-agent — a conversational AI agent that ships near-weekly releases, draws 650+ contributors, and has accumulated over 227k GitHub stars. Model releases continue (NousCoder-14B, Hermes-4.3-36B, nomos-1) but…CompactifAI (Multiverse Computing)CompactifAI (Multiverse Computing)4wMultiverse Computing is executing a deliberate pivot from its quantum-software heritage into a practical AI model-compression and deployment platform under the CompactifAI brand. The evidence depicts an organization scaling enterprise go-to-market across Europe, the Middle East, and North America while investing simultaneously in hard engineering problems: Tensor Network–based LLM compression, VRAM-optimized…Meta AI (Llama)Meta AI (Llama)4wMeta AI is executing a high-stakes organizational pivot from the open-source Llama lineage toward a proprietary, commercially monetized model family — Muse — under the newly formed Meta Superintelligence Labs (MSL). The evidence depicts a lab rebuilding after a self-acknowledged Llama 4 failure, aggressively investing in custom silicon (MTIA spanning four generations), GPU-scale communication frameworks (NCCLX for…Google (DeepMind / Gemini)Google (DeepMind / Gemini)4wGoogle DeepMind is in a period of broad portfolio expansion and high-velocity shipping, but shadowed by a destabilizing dual-leadership departure. The evidence shows a lab simultaneously racing across robotics, materials, bioscience, agents, scientific AI, music generation, and cybersecurity — while the Gemma open-model program drives massive downstream distribution. However, the abrupt exits of CEO Demis Hassabis…Meituan (LongCat)Meituan (LongCat)4wMeituan's LongCat lab has executed one of the fastest documented progressions from modest open-source releases to a 1.6-trillion-parameter, near-frontier agentic coding model — trained entirely on Chinese AI ASIC superpods — in roughly ten months. The lab pursues an explicit model–system co-design philosophy, pairing each model family with custom inference kernels and a growing benchmark portfolio that doubles as…StepFunStepFunJun 27StepFun is a Shanghai-based frontier AI lab executing an unusually broad multimodal strategy—spanning video generation, real-time voice, image editing, 3D asset generation, formal mathematics, GUI agents, and deep research—while releasing the vast majority of its model weights, code, and benchmarks under permissive open-source licenses. The lab's May 2026 release of Step 3.7 Flash, a 198B MoE vision-language model…Xiaomi (MiMo)Xiaomi (MiMo)Jun 27XiaomiMiMo is executing a full-stack, open-source AI strategy that spans text reasoning, vision, audio, embodied AI, and coding agents — with an accelerating pivot toward the Agent era in mid-2026. The lab's evidence trail reveals a deliberate arc: a reasoning-first 7B model family born from pretraining-to-posttraining optimization; rapid horizontal expansion into vision-language (MiMo-VL, May 2025), audio language…Sarvam AISarvam AIJun 27Sarvam AI is executing a three-horizon transition from sovereign AI R&D lab to full-stack platform company competing globally. The evidence pack captures this inflection: a unicorn-level fundraise ($300M at ~$1.5B valuation), the March 2026 release of two MoE reasoning models — Sarvam-30B (32B params) and Sarvam-105B (106B params) — trained from scratch on IndiaAI Mission compute, and a hiring wave of ~50+ open…Upstage (Solar)Upstage (Solar)Jun 27Upstage is executing a platform consolidation play, moving from a pure models-and-APIs company toward an integrated AI stack combining proprietary Solar LLMs, acquired portal (Daum) assets, and agentic platforms (Timely). Core thesis: Upstage is building an AI-for-everyone ecosystem — not just enterprise AI, but consumer-facing AI through a portal. The Solar model family (10.7B→31B→Open 100B→22B in preview/pro)…Reka AIReka AIJun 27Reka is evolving from a multimodal model builder into a physical-AI company. The evidence shows a lab that ships compact, deployment-flexible vision-language models (Flash at 21B, Edge at 7B), layers enterprise agentic platforms on top (Nexus, Vision, Research), and is now orienting research toward world models, embodied data, and physical reasoning. The $110M raise backed by NVIDIA and Snowflake, the Moonvalley…Eigen AIEigen AIJun 27Eigen AI is a 2025-founded inference optimization company acquired by NASDAQ-listed neocloud Nebius for $643M, with the deal closing on 10 June 2026. The lab's optimization stack is being integrated into Nebius Token Factory to deliver production inference at scale. Active hiring for post-training/inference engineering and platform product management signals a dual buildout: continued technical R&D on model serving…WaferWaferJun 27Wafer is a hardware-centric AI inference platform building competitive advantage through GPU kernel optimization expertise, with a distinctive multi-vendor strategy spanning NVIDIA and AMD accelerators. The evidence depicts a company vertically integrated from low-level kernel engineering up to a serverless inference product, using public benchmarks and developer education content as both recruiting and go-to-market…Public AIPublic AIJun 27Public AI is not a frontier model lab; it is a neocloud-adjacent public infrastructure play building an inference utility positioned as an open, democratically governed alternative to commercial AI APIs. The org's GitHub activity reveals a concentrated push to operationalize a chat-based inference platform (chat.publicai.co) atop OpenWebUI, with CI/CD pipeline maturation, API gateway deployment, and multi-geography…Blackbox AIBlackbox AIJun 27Blackbox AI is not a frontier model builder; it is an inference-infrastructure and agent-orchestration platform that competes on serving others' models faster, cheaper, and more securely than anyone else. The company's public signals converge on a single bet: that enterprise and government adoption of coding agents will be won at the orchestration and inference layer, not at the model-training layer [W1, W5]. With…MakoraMakoraJun 27Makora is a performance-engineering organization focused on automated GPU kernel generation and inference optimization. Its public surface spans four categories: (1) an AI-driven kernel generation system (MakoraGenerate) that produces optimized GPU kernels targeting NVIDIA H100/B200, AMD MI300X, and Tenstorrent hardware; (2) a lightweight multi-vendor GPU querying utility (gpuq) supporting CUDA and HIP runtimes; (3)…StreamLake (Kuaishou)StreamLake (Kuaishou)Jun 27Kuaishou's StreamLake is executing a dual-pronged open-source strategy: a video-native multimodal foundation model line (Keye-VL-2.0) aimed at long-video understanding and agentic capabilities, and an AI coding product suite (KAT-Coder, CodeFlicker, Vanchin) targeting the software engineering tool market. Both tracks are anchored in Apache 2.0 releases, rigorous public benchmarking against frontier models (GPT-5,…ParasailParasailJun 27Parasail is an early-stage AI infrastructure company (Series A, $32M raised, $160M valuation) building a serverless inference cloud that aggregates distributed GPU supply into an OpenAI-compatible platform for open-weight models. The company positions itself as a GPU-network orchestration layer: workloads are automatically matched across a multi-provider GPU network, freeing developers from vendor lock-in and…GMI CloudGMI CloudJun 27GMI Cloud is an inference-optimized neocloud building a full-stack platform tightly coupled to NVIDIA's hardware roadmap. The evidence shows a company transitioning from bare-metal GPU provisioning to a managed platform layer: 10 open roles are clustered around a named "Inference Engine" product, AgentBox has shipped as an agent marketplace and hosting platform, and every public post ties GMI's infrastructure…ScalewayScalewayJun 27Scaleway is executing a three-pillar strategy to differentiate as Europe's sovereign AI cloud: (1) serving frontier open-weight models through its Generative APIs platform as a managed alternative to proprietary hyperscalers, (2) embedding sustainability and CSRD compliance tooling directly into its cloud product portfolio, and (3) investing in European AI infrastructure sovereignty through consortia spanning…Novita AINovita AIJun 27Novita AI is executing a two-pronged evolution: it operates a commercial model API and agent sandbox platform for third-party frontier models, while simultaneously building deep inference infrastructure—most visibly pegaflow, a Rust-based KV cache storage engine with vLLM integration—that targets the performance bottleneck of large-scale LLM serving. The pattern of GTM hiring in San Mateo alongside a relentless…SiliconFlowSiliconFlowJun 27SiliconFlow is building a "Token Factory" — an AI inference infrastructure layer that normalizes heterogeneous compute into standardized token output. The GitHub evidence reveals a two-track product strategy: (1) a deep acceleration stack (OneDiff, Nexfort) that compiles diffusion and LLM workloads for faster inference, and (2) BizyAir, a cloud-hosted ComfyUI service that wraps model access into a managed developer…HyperbolicHyperbolicJun 27Hyperbolic is in a post–Series A scaling sprint, pivoting from its early Web3/decentralized microservices roots into a full-stack GPU marketplace aggregator. The evidence shows a company simultaneously hiring for infrastructure depth (GPU orchestration, bare-metal provisioning, SRE), commercial operations (supply, finance, GTM), and developer tooling (CLI, MCP, AI SDK, Gradio). The recent Forge launch crystallizes…FriendliAIFriendliAIJun 27FriendliAI is an AI inference infrastructure company entering an aggressive commercialization phase, signaled by a $20M funding round, a rapid SDK iteration cadence with breaking API changes across all serving tiers, the launch of a public OpenAPI schema, and day-zero support for frontier open-weight models. The dual-hub (Seoul/San Francisco) hiring pattern reveals simultaneous investment in core inference engine…DigitalOcean (GradientAI)DigitalOcean (GradientAI)Jun 27DigitalOcean (GradientAI) is executing a concentrated pivot into agentic AI infrastructure as a managed cloud service, building the full stack from GPU inference to hosted agent runtimes. The evidence reveals a coordinated three-pronged buildout: (1) an Inference Engine now generally available with frontier model support across OpenAI, Anthropic, and fal; (2) a Codex plugin in Public Preview that provisions…Lightning AILightning AIJun 27Lightning AI is in the midst of a structural transformation from developer-framework shop into a vertically integrated neocloud. The merger with Voltage Park [P2, P3] has reshaped the company's operational DNA: it now owns and operates physical data centers across at least three US geographies (Quincy WA, Fort Worth TX, Lisle IL) [P22, P23], is building bare-metal GPU compute, storage, and observability…ReplicateReplicateJun 27Replicate is a post-acquisition platform operating as an inference API aggregator, not a model builder. Following its acquisition by Cloudflare, its activity centers on platform engineering — evidenced by an intense cog release cadence across v0.16–v0.21 — ecosystem integration (SDKs in Python and JavaScript, MCP, LangChain, agent skills), and positioning as the hosted API layer for third-party frontier and…DeepInfraDeepInfraJun 27DeepInfra is an inference-cloud provider exploiting the open-weight model boom, not a model-building lab. Its GitHub footprint reveals a company systematically forking and maintaining the full inference-serving stack — from CUDA kernels to serving engines to client SDKs — while its $107M Series B and targeted hiring confirm a bet on inference infrastructure as a standalone business. The org tracks frontier…Snowflake (Arctic)Snowflake (Arctic)Jun 27Snowflake is executing a deliberate convergence play: its Arctic model family — specialized for SQL, code generation, and enterprise retrieval — is being positioned not as a standalone frontier contender but as the AI inference layer inside a governed, agentic data platform. The firm's public writing, hiring, and releases all orbit a single narrative: "the agentic enterprise". Arctic now spans speculators…SambaNova SystemsSambaNova SystemsJun 27SambaNova Systems is executing a decisive pivot from AI training hardware toward becoming an inference cloud provider purpose-built for agentic AI workloads. The evidence pack captures a company compressing its stack around three interlocking bets: (1) disaggregated/hybrid inference pairing its own SN40 RDU with NVIDIA GPUs for prefill-decode splitting [E29, E53, W2]; (2) "premium inference" as a differentiated…Fireworks AIFireworks AIJun 27Fireworks AI is a Series C ($4B valuation) generative AI infrastructure platform transitioning from inference-speed leader to full-stack AI cloud provider, with training, fine-tuning, serverless and dedicated inference, multi-LoRA serving, and agentic orchestration all built on proprietary infrastructure. The most recent evidence — spanning June 2026 — reveals a company in an intensive GTM buildout phase, anchored…GroqGroqJun 27Groq is rebuilding as a pure-play AI inference cloud after a transformative non-acquisition by Nvidia that took its founding CEO, president, and key engineers. A $650M raise in June 2026 aims to scale GroqCloud to 200MW by 2027 and serve 5M developers on its purpose-built LPU chip architecture. The evidence pack shows Groq rapidly maturing SDK tooling (Python v1.5.0, TypeScript v1.3.0), building an evaluation and…Together AITogether AIJun 27Together AI is consolidating its positioning as the AI-native cloud — an inference-first infrastructure platform that competes on raw speed and cost per token. The evidence pack shows the company simultaneously building out in three directions: (1) deepening the infrastructure surface from GPU clusters into managed storage, networking, and observability, (2) layering enterprise trust and access-control primitives…NebiusNebiusJun 27Nebius is executing a multi-front AI cloud scaling thesis: it is simultaneously building out physical data center capacity across the US and Europe, expanding its GPU orchestration software stack, commercializing a new agentic search product (Tavily), and deepening its research bench through an acqui-hire (Clarifai). The hiring pattern reveals a company transitioning from infrastructure provider to full-stack AI…Zhipu AI (GLM)Zhipu AI (GLM)Jun 8Zhipu AI (GLM) is shipping a broad, fast-moving family of open-weight GLM models across text, vision, OCR, speech, and image generation, releasing point versions at high cadence (GLM-4.5 through GLM-5/5.1 plus specialized variants) and backing them with first-party SDKs. The standout signal is reach: its OCR and Flash models are pulling millions of monthly Hugging Face downloads, and its older ChatGLM line remains…xAIxAIJun 8xAI is in a productize-and-distribute phase: its frontier work (Grok) lives behind the API while the public footprint is dominated by developer tooling and open-weight artifacts of prior generations. The xai-sdk-python is shipping rapidly (five releases tracked, through v1.15.0), and the company has open-sourced both the Grok-1/Grok-2 weights and, notably, the X recommendation algorithm — signaling tight integration…Tencent HunyuanTencent HunyuanJun 8Tencent Hunyuan is running broad on open-weight generative media — its public footprint skews heavily toward 3D, video, image, and world-model generation rather than chat LLMs. The most-downloaded asset is a frontier-scale ~298B-param model (tencent/Hy3-preview, 90k downloads/30d), but the highest-starred surface area on GitHub is its visual-generation stack (Hunyuan3D, HunyuanVideo). It is also pushing into newer…Qwen (Alibaba Cloud)Qwen (Alibaba Cloud)Jun 8Qwen (Alibaba Cloud) is running one of the most prolific open-weight release cadences in the field, shipping a full ladder of dense and Mixture-of-Experts models — currently the Qwen3.5 and Qwen3.6 generations — across every modality and a parallel agentic coding stack (qwen-code, 25k stars). Adoption is enormous: its current flagship-tier checkpoints each pull millions of Hugging Face downloads in a 30-day window.…OpenAIOpenAIJun 8OpenAI is operating on two fronts at once: a frontier-model release cadence aimed at consumers and developers, and a hard pivot into agentic developer tooling. Its public footprint right now is dominated by Codex, a terminal coding agent shipping near-daily alpha builds, and a wave of GPT-5.x launches (GPT-5.5, GPT-5.4, GPT-5.3-Codex) that top Hacker News. The hiring and infra signals point to scaling compute and…NVIDIANVIDIAJun 8NVIDIA is positioning itself as the full-stack supplier of the "AI factory" era — selling not just silicon but open models, agent runtimes, and physical-AI foundation models that run on its hardware. The current push centers on three fronts: long-running agents (the Nemotron 3 Ultra family and the NemoClaw agent blueprint), physical/world AI (Cosmos 3 and robotics), and local/personal agents on new hardware (RTX…Moonshot AI (Kimi)Moonshot AI (Kimi)Jun 8Moonshot AI (Kimi) is shipping open-weight, trillion-parameter mixture-of-experts frontier models at a fast iteration cadence — the Kimi-K2 line is its flagship, now through K2.5 and K2.6 plus a dedicated K2-Thinking variant. Alongside the weights it is building a full agentic-coding surface (the kimi-cli / kimi-code tools) and publishing efficiency-oriented architecture research (linear attention, attention…Mistral AIMistral AIJun 8Mistral AI is executing a broad open-weights strategy across every modality and size tier at once: text instruct/reasoning models from 3B up to 128B, a Voxtral audio/speech family (realtime, TTS), and a Devstral coding line. Distribution runs through Hugging Face at serious volume and a full client/tooling stack (mistral-inference, mistral-common, multi-language SDKs). The 2512/2602/2603 release cadence shows rapid,…