Neocloudfresh 4d

Cerebras

Signal timeline329 total
Jun 23, 2026
4wJobHardware / Low Level Security EngineerRemote, California, United States; Sunnyvale CA or Toronto Canadasource ↗
4wJobNetwork Security EngineerRemote, California, United States; Sunnyvale CA or Toronto Canadasource ↗
4wJobPrincipal Network Security ArchitectRemote, California, United States; Sunnyvale CA or Toronto Canadasource ↗
4wJobPrincipal AI Security EngineerRemote, California, United States; Sunnyvale CA or Toronto Canadasource ↗
Jun 20, 2026
Jun 19, 2026
Jun 16, 2026
Jun 16JobML Systems Performance EngineerBengaluru, Karnataka, Indiasource ↗
Jun 12, 2026
Jun 12JobNetwork Engineer Sunnyvale, CA; Toronto, Ontario, Canadasource ↗
Jun 10, 2026
Jun 10JobApplied Machine Learning Research ScientistSunnyvale CA or Toronto Canadasource ↗
Jun 10JobProduct Manager, Strategic Verticals San Francisco, California, United Statessource ↗
Jun 10JobLead Full Stack Machine Learning EngineerBengaluru, Karnataka, Indiasource ↗
Jun 9, 2026
Jun 9JobSenior / Staff Technical Program Manager - Datacenter Capacity Delivery (E2E) Europe; Remote, California, United States; Sunnyvale, CA; Toronto, Ontario, Canadasource ↗
Jun 8, 2026
Jun 8JobDesign Verification EngineerBengaluru, Karnataka, Indiasource ↗
Jun 5, 2026
Jun 4, 2026
Jun 4JobASIC ArchitectSunnyvale, CAsource ↗
Jun 3, 2026

Top signals

  1. #1WritingCerebras Launches Qwen3 235b World S Fastest Frontier Ai Model With Full 131k Context Support10.0
  2. #2WritingCerebras Launches Worlds Fastest Deepseek R1 Llama 70b Inference9.0
  3. #3WritingCerebras Systems Closes Usd850 Million Revolving Credit Facility9.0
  4. #4WritingIntroducing Condor Galaxy 1 A 4 Exaflop Supercomputer For Generative Ai9.0
  5. #5WritingCepo Update Turbocharging Reasoning Models Capability Using Test Time Planning8.0

Agent answer

Cerebras has 329 loaded public signals: 107 hiring, 8 forks, 38 releases or model cards, 160 talking, and 16 repos. Latest signal: Cerebras/capi-image-builder. Data-business radar is currently scoped to frontier labs, so this category does not expose radar lanes. The standing analysis was generated with deepseek-v4-pro and 92 evidence refs.

Cerebras

has loaded 329 public signals

Cerebras

has hiring signal count 107

Cerebras

has fork signal count 8

Cerebras

has release signal count 38

Analysis — agent synthesisfull report →generated June 30, 2026

Thesis

Cerebras is executing a hard pivot from wafer-scale training hardware specialist to full-stack inference cloud provider. The evidence shows a company that IPO'd in May 2026 E58, closed an $850M revolving credit facility E57, and is aggressively building out inference datacenters across North America and Europe with a target of 20x aggregate capacity expansion P22. The wafer-scale architecture delivers inference speeds 10–20x faster than GPU clouds for frontier models like Llama 3.1 405B and Kimi K2.6 P21P25W1, and Cerebras is now landing anchor tenants including OpenAI (750 MW partnership), Mistral, Perplexity, and HuggingFace P22W2. Hiring signals confirm the inference-first strategy, with concentrated openings in inference platform engineering, datacenter operations, and a new security engineering cluster — all while silicon design hiring continues in Sunnyvale, Bengaluru, and Toronto . The public narrative has shifted from "training made easy" to "instant-speed inference" and sovereign AI infrastructure E59P23.

Signal desks

Hiring

  • Inference platform is the #1 hiring priority. Roles span Software Engineer through Staff Software Engineer for Inference Platform in Sunnyvale E9E10, Staff Inference ML Runtime Engineer E43, Senior Performance Engineer – Inference E41, and LLM Inference Performance & Evals Engineer in Toronto E35. A Staff Software Engineer posting describes a "next-generation architecture of a globally distributed inference platform" W4.
  • Security engineering surge (June 2026). Four security roles opened within days: Hardware / Low Level Security Engineer E3, Network Security Engineer E4, Principal Network Security Architect E5, and Principal AI Security Engineer E6 — all listing Remote CA, Sunnyvale, or Toronto. This cluster follows a blog on AI inference cybersecurity E12 and likely signals hardening for enterprise and sovereign cloud deployments.
  • Datacenter buildout hiring. Head of Data Center Acquisition E51, Director / Senior Director of Critical Facility Operations E34, Business Operations Lead – Datacenters E31, and Senior/Staff TPM – Datacenter Capacity Delivery (E2E, with a Europe location option) E22 — all directly support the six-new-datacenters expansion plan P22.
  • Silicon design continues at scale. ASIC Architect E32, Senior Front End Design Engineer (Microarchitecture) in both Bengaluru E8 and Sunnyvale E23, multiple Design Verification Engineers at various levels E24E29E30, Physical Design Engineer E19, 3D Physical Design Engineer E28, Post Silicon Bring-Up E50, and Design Validation Test Lead E46. This signals ongoing investment in next-generation wafer-scale silicon.
  • Kernel and compiler engineering. Multiple Kernel Engineers across Sunnyvale, Toronto, Bengaluru, and Remote CA E44E49E52, plus Advanced Technology Compiler Engineer (Sunnyvale, Vancouver) E37 — sustained investment in the software stack that maps ML workloads to the wafer-scale fabric.
  • ML research and applied engineering. Applied Machine Learning Research Scientist E18, Lead Full Stack ML Engineer – Bengaluru E21, ML Systems Performance Engineers in Bengaluru and Toronto E16E55, Senior ML Software Engineer – Integration & Quality E40, and ML Software Tool Development Engineer E48.
  • GTM and commercial scaling. Vice President of Creative & Integrated Marketing E45, Product Manager – Strategic Verticals (San Francisco) E20, and Sr. Sourcing Manager – Critical Components E11 point to supply chain and go-to-market maturation.
  • Geographic hubs. Sunnyvale remains the primary hub. Toronto is now a major secondary center for inference, ML, and security. Bengaluru is growing for silicon design, verification, kernel, and full-stack ML. Vancouver appears for compiler roles E37. European datacenter delivery roles signal physical expansion outside North America E22.

Forks

No cited evidence in this pack. The single repository found — Cerebras/ghcp-mcp-registry-test P1E1 — is an original repo (not a fork), created June 2026 with 0 stars, described as a "GHCP MCP Registry Test," suggesting internal tooling or Model Context Protocol experimentation rather than upstream dependency tracking.

Releases

  • Cerebras/sdk-examples v2.10.0 — released April 21, 2026 E60. Indicates ongoing SDK iteration but no release notes or artifacts are captured in this evidence pack to assess what changed.
  • No additional model cards, versioned packages, or repository releases with inspectable artifacts are cited in this evidence. Shipping is primarily signaled through blog posts and press announcements rather than GitHub releases (see Shipping section below).

Talking

  • Inference speed as the core narrative. Cerebras's public output is dominated by inference speed benchmarks: Llama 3.1 405B at 969 tok/s, 12x faster than GPU clouds P21; Llama 3.1 70B breaking 2,100 tok/s, 16x faster than the fastest GPU solution P25; Kimi K2.6 (trillion-parameter) at nearly 1,000 tok/s for enterprise E25W1; head-to-head comparison "Which Is Faster: Gemini 3.5 Flash or Kimi K2.6 on Cerebras" E14; and the Gemma 4 multimodal inference launch E13.
  • Institutional/financial milestones. IPO announcement E58 and $850M credit facility close E57 both published late May 2026. VentureBeat coverage frames the IPO as "the largest tech IPO of 2026" W1.
  • Sovereign AI positioning. "What Is Sovereign AI and How Cerebras Helps Nations" E59 signals a government/national-infrastructure GTM motion, leveraging the datacenter buildout and G42 partnership.
  • Research thought leadership. Posts on agent verification ("Never Loop Without Verifiers") E2, the economics of AI reasoning E15, MoE architectures ("MoE Guide Calculator") E7, AI inference cybersecurity E12, and the speed-accuracy tradeoff E56 position Cerebras as a research-aware infrastructure provider.
  • Developer and product content. "Building an AI-Powered Search Assistant for Zoom Team Chat" P24, "Chatting Your Way Through 4500 NeurIPS Papers" P26, and "Generating Beautiful UIs" E27 demonstrate product integrations and developer tooling.
  • Traction is thin on community channels. The Kimi K2 Enterprise post registered only 2 points and 0 comments on HN E25. Third-party validation comes primarily from VentureBeat W1 and Artificial Analysis benchmarks cited in Cerebras's own posts P21P25.

Shipping

  • Cerebras Inference API is live with Llama 3.1 8B, 70B, and 405B at production scale. The 405B deployment achieves 969 tok/s output and 240ms time-to-first-token at 128K context length P21. Llama 3.1 70B runs at 2,100 tok/s after a 3x software-only performance upgrade P25.
  • Kimi K2.6 — a trillion-parameter open-weight model from Moonshot AI — is now served to enterprise customers at ~1,000 tok/s, which VentureBeat reports is nearly 7x faster than GPU clouds W1E25.
  • Gemma 4 multimodal inference launched June 2026, described as "the fastest inference is now multimodal" E13.
  • Six new AI inference datacenters announced for 2025: Santa Clara, Stockton, Dallas (online), Minneapolis (Q2 2025), Oklahoma City and Montreal (Q3 2025), plus Midwest/Eastern US and Europe (Q4 2025). Aggregate capacity expanding 20x, with 85% of capacity in the United States P22.
  • Cerebras/sdk-examples v2.10.0 shipped April 2026 E60.
  • Enterprise customer logos: Mistral (Le Chat assistant), Perplexity (AI search), HuggingFace, and AlphaSense are cited as inference customers P22. OpenAI partnership at 750 MW scale cited in job postings W2W3.
  • Historical shipping of note: BTLM-3B-8K trained on Condor Galaxy 1 with Opentensor P10; Jais 13B Arabic LLM with G42's Inception P19; CSoft R1.3 enabling GPT-J training P7.

Research themes

  • Sparse pre-training for LLMs. Published SPDF (Sparse Pre-training and Dense Fine-tuning) at ICLR 2023 workshop, demonstrating up to 75% unstructured sparsity on GPT-3 XL with 60% fewer training FLOPs while preserving downstream accuracy P9.
  • Variable sequence length (VSL) training. Method to reduce wall-clock time for long-context LLM training by 29% FLOPs without architecture changes, using staged sequence length curriculum (2K→8K) P8.
  • bfloat16 / automatic mixed precision. Published analysis of bfloat16 benefits for GPT-style model training on wafer-scale hardware P12.
  • High-resolution computer vision on wafer-scale. Training on 50-megapixel images with the CS-2, exploiting on-chip memory to bypass GPU memory bottlenecks for image segmentation and diffusion models P11.
  • Mixture of Experts (MoE). Published an "MoE Guide Calculator" E7, signaling active work on sparse expert architectures.
  • AI reasoning economics. Blog analyzing the cost/speed tradeoffs of reasoning models E15, connecting directly to Cerebras's inference speed advantage.
  • Agent verification. "Never Loop Without Verifiers" E2 addresses agentic AI safety and reliability — a research theme aligned with enterprise deployment concerns.
  • AI inference cybersecurity. Published post on securing inference infrastructure E12, paired with the security hiring cluster .

Hiring & scaling

Cerebras is hiring across five distinct vectors, with inference platform and datacenter operations receiving the heaviest investment:

1. Inference platform engineering (Sunnyvale, Toronto): Software and Staff Engineers building a globally distributed inference platform E9E10E43E41E35W4. 2. Datacenter operations and delivery (Sunnyvale, Europe, Remote): From Head of Data Center Acquisition through Critical Facility Operations Director and capacity delivery TPMs E51E34E31E22. 3. Silicon design and verification (Sunnyvale, Bengaluru): ASIC Architect, Front End Design, Physical Design (including 3D), Design Verification at all levels E32E8E23E19E28E24E29E30E50E46. 4. Security engineering (Remote CA, Sunnyvale, Toronto): A concentrated cluster of four security roles spanning hardware, network, and AI security E3E4E5E6. 5. ML research and kernel/compiler (Sunnyvale, Toronto, Bengaluru, Vancouver, Remote): Applied ML Scientists, Kernel Engineers, Compiler Engineers, ML Performance Engineers E18E21E44E49E52E37E16E55E40E48.

Geographic expansion is material: Bengaluru is becoming a silicon and ML engineering hub; Toronto is the inference and security co-headquarters; Vancouver appears for compiler; European roles target datacenter delivery E22. GTM hiring (VP Marketing, Product Manager for Strategic Verticals) indicates commercialization maturation E45E20.

Category implications

  • Infrastructure strategy: inference-first cloud with wafer-scale hardware as the moat. Cerebras's wafer-scale engine eliminates the memory bandwidth bottleneck that forces GPUs to move model weights from off-chip HBM on every token generation pass P23. The result is inference speeds that GPU clouds cannot match without architectural change. The six-datacenter buildout P22, 20x capacity expansion, and 85% US capacity allocation signal an intention to become the default high-speed inference layer for frontier models — not just a hardware supplier.
  • Product: from training systems to inference-as-a-service. Early evidence (2021–2023) shows Cerebras positioned as a training accelerator for pharma, finance, and research institutions P4P13P14P16. The current evidence (2024–2026) shows a completed pivot to inference API and cloud, with training mentioned only historically. The SDK examples release E60 and PyTorch integration P27P28 provide developer on-ramps, but the core product is now instant-speed inference.
  • Research: sparse training, long context, and MoE align with inference cost reduction. Cerebras's published research on sparse pre-training P9, VSL P8, and MoE E7 maps directly to reducing the compute cost of serving large models — which is the core value proposition of their inference cloud. Agent verification research E2 and reasoning economics E15 extend the narrative into enterprise reliability.
  • Hiring: security cluster signals enterprise/sovereign readiness. The four simultaneous security engineering openings in June 2026 , combined with the AI inference cybersecurity post E12 and sovereign AI positioning E59, suggest Cerebras is hardening its inference cloud for government, defense, and regulated enterprise workloads — a necessary step for the sovereign AI GTM motion.
  • GTM: OpenAI anchor tenant plus sovereign and enterprise verticals. The OpenAI 750 MW partnership W2W3 provides a demand anchor that justifies the datacenter CapEx. Mistral, Perplexity, HuggingFace, and AlphaSense add commercial logos P22. Sovereign AI messaging E59 and the G42/Condor Galaxy partnership P10P19 add government and Middle East dimensions. The VP of Creative & Integrated Marketing hire E45 and Product Manager for Strategic Verticals E20 indicate the GTM organization is being built out to match the infrastructure investment.
  • Caution: thin third-party validation outside self-reported benchmarks. Community traction is low (2 HN points on the Kimi K2 Enterprise post E25). VentureBeat coverage W1 and Artificial Analysis benchmarks (cited in Cerebras's own posts P21P25) are the primary external validation. Customer case studies in the evidence pack date from 2021–2022 and describe on-premise CS-1/CS-2 deployments P13P14, not the current inference cloud product.

Traction highlights

  • IPO completed — described by VentureBeat as the largest tech IPO of 2026 E58W1.
  • $850M revolving credit facility closed May 2026 E57.
  • OpenAI multi-year partnership at 750 MW scale, cited across multiple job postings W2W3.
  • Inference performance records: Llama 3.1 405B at 969 tok/s (75x faster than AWS) P21; Llama 3.1 70B at 2,100 tok/s (16x faster than fastest GPU) P25; Kimi K2.6 trillion-parameter model at ~1,000 tok/s W1.
  • Named inference customers: Mistral (Le Chat), Perplexity, HuggingFace, AlphaSense P22.
  • Six new datacenters adding 20x capacity, 85% US-based P22.
  • Historical model deliveries: BTLM-3B-8K with Opentensor on Condor Galaxy 1 P10; Jais 13B Arabic LLM with G42 P19; GSK epigenomic models with 15x training speedup P13; financial services BERT-LARGE at 15x speedup vs. 8-GPU server P14.
Deep reports