Neocloudfresh 1w

Cerebras

Signal timeline338 total
Feb 25, 2026
Feb 25Modelcerebras/Step-3.5-Flash-REAP-121B-A11BNew model release, low traction.sourcenotability 5.0/1013615
Feb 25Modelcerebras/Step-3.5-Flash-REAP-149B-A11BNotable model, minimal tractionsourcenotability 5.0/101075
Jan 23, 2026
Jan 23Modelcerebras/GLM-4.7-Flash-REAP-23B-A3BLow traction model releasesourcenotability 3.0/1091080
Jan 10, 2026
Jan 10Modelcerebras/GLM-4.7-REAP-218B-A32BLow traction (48 downloads), niche modelsourcenotability 3.0/1013833
Jan 10Modelcerebras/GLM-4.7-REAP-268B-A32BLow downloads, minor model releasesourcenotability 2.0/1010419
Dec 9, 2025
Dec 9Modelcerebras/DeepSeek-V3.2-REAP-345B-A37BFine-tune of DeepSeek V3, moderate downloadssourcenotability 5.0/1013934
Nov 6, 2025
Nov 6Modelcerebras/Kimi-Linear-REAP-35B-A3B-InstructLow traction model releasesourcenotability 4.0/1026570
Oct 30, 2025
Oct 30Modelcerebras/Qwen3-Coder-REAP-363B-A35BLarge model variant but low tractionsourcenotability 5.0/10835
Oct 30Modelcerebras/Qwen3-Coder-REAP-246B-A35B16 downloads in 30 days, very low tractionsourcenotability 1.0/10828
Oct 24, 2025
Oct 24Modelcerebras/GLM-4.6-REAP-268B-A32BVery low downloads (17), likely minorsourcenotability 3.0/109912
Oct 24Modelcerebras/GLM-4.6-REAP-252B-A32BLow traction releases, minor despite size.sourcenotability 3.0/10874
Oct 23, 2025
Oct 23Modelcerebras/GLM-4.6-REAP-218B-A32BVery low traction (14 downloads), routine releasesourcenotability 2.0/109417
Oct 20, 2025
Oct 20Modelcerebras/GLM-4.5-Air-REAP-82B-A12BLow traction, niche variantsourcenotability 1.0/10238115
Oct 20Modelcerebras/Qwen3-Coder-REAP-25B-A3BNotable release from Cerebras but low traction.sourcenotability 5.0/1066589
Mar 4, 2025
Mar 4Modelcerebras/Llama-3-CBHybridM-8BLow traction, small modelsourcenotability 3.0/1071
Feb 26, 2025
Feb 26Modelcerebras/Llama-3-CBHybridL-8BVery low downloads (8) in 30 dayssourcenotability 1.0/1077
Aug 16, 2024
Aug 16Modelcerebras/Dragon-DocChat-Context-EncoderLow traction, routine model release.sourcenotability 3.0/101031
Aug 16Modelcerebras/Dragon-DocChat-Query-Encoder6 downloads in 30 days, negligible tractionsourcenotability 1.0/10853
Aug 13, 2024
Aug 13Modelcerebras/Llama3-DocChat-1.0-8BVery low traction, minor fine-tunesourcenotability 1.0/107869
Mar 22, 2024
Mar 22Modelcerebras/Cerebras-GPT-IntermediateIntermediate-sized model release from Cerebras, not a flagship.sourcenotability 5.0/10
Mar 19, 2024
Mar 19Modelcerebras/Cerebras-ViT-L-336-patch14-llava13b-ShareGPT4VLow-traction VL model release.sourcenotability 2.0/1065
Mar 19Modelcerebras/Cerebras-LLaVA-13BLow traction model releasesourcenotability 3.0/10756
Mar 19Modelcerebras/Cerebras-ViT-L-336-patch14-llava7b-ShareGPT4VLow traction model release, 9 downloads in 30 days.sourcenotability 3.0/10611
Mar 19Modelcerebras/Cerebras-LLaVA-7BLow traction LLaVA fine-tune, only 11 downloads.sourcenotability 3.0/10712
Dec 5, 2023
Dec 5Modelcerebras/btlm-3b-8k-chatSmall model, negligible downloads.sourcenotability 3.0/10914

Top signals

  1. #1WritingCerebras Launches Qwen3 235b World S Fastest Frontier Ai Model With Full 131k Context Support10.0
  2. #2WritingCerebras Launches Worlds Fastest Deepseek R1 Llama 70b Inference9.0
  3. #3WritingCerebras Systems Closes Usd850 Million Revolving Credit Facility9.0
  4. #4WritingIntroducing Condor Galaxy 1 A 4 Exaflop Supercomputer For Generative Ai9.0
  5. #5WritingCepo Update Turbocharging Reasoning Models Capability Using Test Time Planning8.0

Agent answer

Cerebras has 338 loaded public signals: 107 hiring, 8 forks, 38 releases or model cards, 169 talking, and 16 repos. Latest signal: How Cerebras Serves Gpt 5 6 Sol At Up To 750 Tokens Per Second. Data-business radar is currently scoped to frontier labs, so this category does not expose radar lanes. The standing analysis was generated with deepseek-v4-pro and 94 evidence refs.

Cerebras

has loaded 338 public signals

Cerebras

has hiring signal count 107

Cerebras

has fork signal count 8

Cerebras

has release signal count 38

Analysis — agent synthesisfull report →generated September 4, 2026

Thesis

Cerebras is consolidating into a speed layer for frontier inference, not a model house. The through-line across the pack is that wafer-scale hardware removes the GPU memory wall and the multi-chip networking tax, letting Cerebras serve the same frontier model — GPT-5.6 Sol with the same architecture, weights, precision, context configuration, and reasoning settings — at up to 750 output tokens/second, where "the only difference is the hardware" P1P5W1. That speed-to-intelligence argument is the public framing ("speed and intelligence are no longer mutually exclusive"; "in AI, speed is productivity") W1W5.

The operational signals point to a two-sided buildout: (1) a hardware, manufacturing, and datacenter scale-up — the CS-4/Nexus rack-scale system, Flex's ~7x CS-3 production expansion in Milpitas, and a dense wave of datacenter, security, and silicon roles — and (2) an inference *platform* plus GTM push — a vLLM-disagg fork, AMD disaggregated inference, and partnerships with OpenAI, Lovable, Upstage, and AMD P2P3P9P12P16. Strategically, serving frontier models is framed as giving Cerebras early visibility into where the industry is heading, a position "previously available only to NVIDIA" W2.

Signal desks

Hiring

  • Security cluster: Hardware/Low Level Security Engineer, Network Security Engineer, Principal Network Security Architect, and Principal AI Security Engineer — all in Sunnyvale CA / Toronto / Remote CA — signals hardening of hardware and inference-cloud surfaces E24E25E26E27.
  • Inference platform/runtime: Software Engineer and Staff SWE, Inference Platform (Sunnyvale); Staff Inference ML Runtime Engineer; Senior Performance Engineer, Inference; LLM Inference Performance & Evals Engineer (Toronto); ML Systems Performance Engineer (Bengaluru); Senior ML SW Engineer Integration & Quality — a repeated cluster around inference serving, performance, and evals E29E30E34E52E57E58E60.
  • Silicon/ASIC roadmap: Senior Front End Design (Microarchitecture) in Sunnyvale and Bengaluru; Physical Design; 3D Physical Design; Design Verification (Sunnyvale and Bengaluru); ASIC Architect — consistent with the previewed CS-5/CS-6 roadmap E28E37E41E42E45E46E47E49.
  • Datacenter capacity & ops: TPM Datacenter Capacity Delivery (E2E); Director/Sr Director Critical Facility Operations; Business Operations Lead, Datacenters; Prognostics & Health Monitoring Engineer; Sr. Sourcing Manager – Critical Components E31E40E43E48E51.
  • Locations: Sunnyvale CA dominates; Toronto Canada, Bengaluru India, Vancouver BC, San Francisco, and Europe/Remote also present E24E28E38E40E54.

Forks

  • Cerebras/vllm-disagg — fork of vllm-project/vllm (Python, Apache-2.0, low stars); the "disagg" naming maps directly to the disaggregated-inference strategy announced with AMD P9P10.
  • Cerebras/capi-image-builder — fork of kubernetes-sigs/image-builder for ClusterAPI-compatible Kubernetes VM images; points to datacenter automation / self-managed infrastructure P12E14.
  • Cerebras/ghcp-mcp-registry-test — a new (non-fork) MCP registry test repo; low-profile signal of agent/MCP tooling-registry exploration P18E20.

Releases

  • CS-4 + Nexus: first multi-wafer rack (three WSE-3 Turbo processors), up to 30x faster inference than GPUs, shipping this quarter P2P3W3W5.
  • GPT-5.6 Sol Ultrafast (with OpenAI): up to 750 tok/s, same architecture/weights/precision/context as the standard endpoint, limited preview in the OpenAI API P1P5W1.
  • AMD disaggregated inference: AMD Helios + Cerebras WSE as a single workflow, up to 5x T/s/W, first available through Cerebras Cloud in H2 2026 P9E2.
  • Serving partnerships: Lovable running latency-sensitive workloads on dedicated Cerebras capacity P6; Upstage Solar 31B at up to 2,000 tok/s for Korea P14; Gemma 4 31B multimodal at ~2,300 tok/s P17.
  • Manufacturing: Flex expansion for ~7x CS-3 production in Milpitas, CA P16.
  • Open model artifact: Cerebras-GPT-590M (Apache-2.0, The Pile, Chinchilla scaling) — dated, but the only model card in this pack P11.

Talking

  • Speed-to-intelligence: "speed and intelligence are no longer mutually exclusive" (Feldman) W1; "in AI, speed is productivity," and 30x speed "gives an agentic system room" W5.
  • Hardware explainer: the GPU memory wall and the "only difference is the hardware" framing for Sol Ultrafast P1.
  • Internal tooling: the Cerebras Knowledge base — 15,000 questions/day, embeddings/PGVECTOR retrieval — signals internal agent/RAG adoption P8E9.
  • AI-native hiring narrative: interviews now permit AI tools while evaluating judgment, verification, communication, and ownership P15E15.
  • Research-adjacent writing: the economics of AI reasoning E33; "never loop without verifiers" E23; MoE guides E3E5E11; AI inference cybersecurity E12.
  • Earnings framing: serving frontier models gives early visibility into industry direction W2.

Shipping

Concrete shipped/announced artifacts in this pack:

  • CS-4 is the headline: a rack-scale Nexus platform with three WSE-3 Turbo processors, up to 2x CS-3 and up to 30x GPU tokens-per-second-per-user, shipping this quarter P3W3W5.
  • OpenAI Ultrafast mode launched in limited preview in the OpenAI API, at up to 750 tok/s P5W1.
  • AMD disaggregated inference unveiled at Advancing AI 2026, expected H2 2026 via Cerebras Cloud P9.
  • Gemma 4 31B hosted at ~2,300 tok/s and Upstage Solar 31B at up to 2,000 tok/s live on Cerebras Inference Cloud P14P17.
  • Flex Milpitas lines to scale CS-3 production ~7x through 2026 P16.
  • Older artifacts (CSoft R1.3, BTLM-3B-8K, Cerebras-GPT) show a long tail of training-tooling/open-model releases, but they are dated relative to the current inference push P11P24P27.

Most "release" evidence is press/blog framing; the only inspectable code/model artifacts are the vLLM-disagg fork, the capi-image-builder fork, and the Cerebras-GPT-590M card — the pack is thin on actual public repos/cards for the 2026 launches P10P11P12.

Research themes

  • Sparse attention for KV-cache reduction: hybrid dense/sparse Llama-3.1-8B-Instruct released with ~50% KV cache memory reduction while largely maintaining long-context performance W4.
  • Sparse pre-training + dense fine-tuning (SPDF): pre-train GPT-3 XL with up to 75% unstructured sparsity and 60% fewer FLOPs, then dense fine-tune to preserve accuracy P26.
  • Long-context training efficiency: Variable Sequence Length (VSL) reduces FLOPs ~29% vs fixed 8k training P25.
  • Scaling laws: Cerebras-GPT family trained Chinchilla-optimal (20 tokens/param) on the Andromeda CS-2 supercomputer P11.
  • Multimodal serving: Gemma 4 vision/reasoning/long-context/function-calling at interactive speeds P17.
  • Applied internal RAG: distillation → embeddings (PGVector 3072-dim) → retrieval → fusion/rerank → synthesis P8.

The Cerebras-authored research in this pack skews to training-efficiency methods (sparsity, VSL) and serving/multimodal integration; there is no evidence of a Cerebras-trained frontier model in the current window P11P25P26W4.

Hiring & scaling

Hiring clusters reveal four scaling priorities : 1. Silicon roadmap — microarchitecture, physical/3D design, verification, and ASIC architect across Sunnyvale and Bengaluru E28E37E41E42E45E46E47E49. 2. Inference software/platform — inference platform SWEs, ML runtime, performance, and eval roles E29E30E34E52E57E58E60. 3. Datacenter capacity & ops — capacity delivery TPM, critical facility ops, datacenter business ops, critical-component sourcing, and prognostics/health monitoring E31E40E43E48E51. 4. Security — hardware/low-level, network, and AI security (Principal-level) E24E25E26E27.

Geography is broadening from Sunnyvale to Toronto, Bengaluru, Vancouver, San Francisco, and Europe/Remote, consistent with scaling both silicon and cloud operations E24E28E38E40E54. Manufacturing scale is also explicit: Flex Milpitas adds production lines, floor space, test infrastructure, and skilled manufacturing talent for ~7x CS-3 output P16. On the people-process side, Cerebras is standardizing AI-assisted technical interviews across engineering P15.

Category implications

  • Neocloud/modular infrastructure: CS-4 is explicitly pitched to "neoclouds and hyperscalers" needing modular systems "quickly manufactured, installed, expanded, and upgraded at gigawatt scale," with data-center operators optimizing throughput per gigawatt P3P4. The Nexus "compute backpack" modular power/cooling/I/O design and on-site backpack installation is a datacenter-deployment strategy, not just a chip P2.
  • Inference product strategy: disaggregation is the clear product thesis — high-throughput AMD Instinct GPUs plus ultra-low-latency Cerebras WSE in one workflow, up to 5x T/s/W, distributed first through Cerebras Cloud P9. The vLLM-disagg fork is the software-side evidence of this P10.
  • GTM: tiered serving partnerships — OpenAI (frontier-model speed), Lovable (interactive software creation), Upstage (Korea enterprise), AMD (disaggregated infra) — show Cerebras selling into both model providers and application platforms P1P6P9P14.
  • Research/roadmap: the CS-6 preview (3D-stacked DRAM above wafer logic) attacks the memory-capacity problem directly, extending the memory-wall narrative W6; eval-specific hiring (LLM Inference Performance & Evals Engineer) supports benchmarking-led selling E52.
  • Strategic data advantage: serving closed frontier models gives Cerebras early visibility into emerging AI technologies, "previously available only to NVIDIA" W2.
  • Hiring/GTM implication: a Product Manager, Strategic Verticals role plus South Korea/Sweden partnership activity suggest vertical and geographic expansion beyond pure infrastructure E38P6P14.

No vendor revenue or market-share claims are supportable from this pack; the only financial events cited are a $1B Series H raise and the Q2 2026 earnings-transcript framing E22W2.

Traction highlights

  • The OpenAI Ultrafast post drew 176 points / 52 comments on HN — the strongest discussion signal in the pack E1.
  • GPT-5.6 Sol Ultrafast benchmark: HLE's 2,500 questions in 11h11m vs Claude Fable 5's 78h27m (~7x faster at comparable accuracy) P5.
  • Q2 2026: "completed support for OpenAI's GPT-5.6 Sol," served at 10x speed, offered through Cerebras Cloud W2.
  • Lovable: >50M projects built since Nov 2024, now running latency-sensitive workloads on Cerebras P6.
  • Upstage Solar 31B at up to 2,000 tok/s; Gemma 4 31B at ~2,300 tok/s (first open-weight multimodal past 2,000 tok/s on Cerebras) P14P17.
  • CS-4 up to 2x CS-3 and up to 30x GPU tokens-per-second-per-user W3.
  • $1B Series H (Feb 2026) E22.
  • The AMD announcement drew 27 points / 9 comments on HN; other posts (MoE scale, knowledge base, Kimi K2) show minimal HN traction E2E3E9E21.
Deep reports