Neocloudfresh 2h

Baseten

Signal timeline583 total
Sep 8, 2026
Sep 5, 2026
Sep 3, 2026
Sep 2, 2026
Sep 1, 2026
Aug 29, 2026
Aug 27, 2026
1wJobProduct DesignerSan Franciscosource ↗
Aug 25, 2026
Aug 24, 2026
Aug 21, 2026
Aug 20, 2026
2wJobAI EngineerSan Franciscosource ↗
Aug 19, 2026
Aug 18, 2026

Top signals

  1. #1WritingDeepseek V4 Pro 0813 Available On Baseten9.0
  2. #2WritingSota Performance For Gpt Oss 120b On Nvidia Gpus8.0
  3. #3WritingHow We Made The Fastest Gpt Oss On Nvidia Gpus 60 Percent Faster7.0
  4. #4WritingIntroducing Nemotron 35 Lightning7.0
  5. #5WritingKimi K2 Explained The 1 Trillion Parameter Model Redefining How To Build Agents7.0

Agent answer

Baseten has 583 loaded public signals: 129 hiring, 68 forks, 100 releases or model cards, 239 talking, and 47 repos. Latest signal: How Baseten Makes Pyannotes Diarization Models 96x Faster. Data-business radar is currently scoped to frontier labs, so this category does not expose radar lanes. The standing analysis was generated with deepseek-v4-pro and 92 evidence refs.

Baseten

has loaded 583 public signals

Baseten

has hiring signal count 129

Baseten

has fork signal count 68

Baseten

has release signal count 100

Analysis — agent synthesisfull report →generated September 9, 2026

Thesis

Baseten is a neocloud inference platform in post-Series-F scale-up mode: nearly every job posting cites a recent $1.5B Series F led by Altimeter Capital, Conviction Partners, and Spark Capital P4P8P9P14P15P16P17P21P22P28. The evidence shows two parallel moats being built at once — (1) a physical GPU capacity/orchestration layer (NVIDIA Blackwell B200 readiness, multi-cloud, on-prem data-center delivery) P4P8P9, and (2) an enterprise-plus-agent-native commercialization layer (fine-grained IAM, finance/GTM/data hires, MCP and Claude Code tooling) P15P16P17P19P22. At the same time it is pushing into post-training/RL infrastructure (managed rollouts, baseten train/loops) and launched a public research arm, Base Labs, with an open-publication mandate P6P18W3W4. Net signal: Baseten is moving from "fast model serving" toward owning the fleet, the training loop, and the enterprise trust surface of the neocloud category.

Signal desks

  • Hiring — capacity/compute is the densest cluster. Global Capacity Manager (Compute) owns the GPU fleet, Kubernetes orchestration, and Blackwell (B200) readiness, including Go-based operators to triage unhealthy H100 nodes P4E4; Delivery Director, Capacity Programs (G&A) runs on-prem data-center and neocloud GPU delivery P8E7; Capacity Operations Manager (G&A) owns fleet health, observability, utilization, and remediation P9E6.
  • Hiring — training/RL and model performance. AI Engineer (Training Platform) E47, Forward Deployed Engineer (Training) E50, and Technical Program Manager, Model Performance E25 are all San Francisco-based.
  • Hiring — enterprise & data. Founding Software Engineer, Identity and Authorization (Enterprise) for fine-grained auth and agentic-workload credentials P15E16; first dedicated Marketing Analytics Manager P14E14, Product Data Scientist P16E13, and Revenue Analyst (dbt, Salesforce/CPQ, Orb metering) P17E12.
  • Hiring — GTM/finance scale-up. Multiple Revenue Strategy & Ops Sr. Analyst roles P22P28E19E24, GTM Systems Manager E30, GTM Engineer (New York) E41, Sales Development Manager and inbound SDR E42E51, Product Marketing Manager, Model API E48, Finance Systems Lead P21E22, Revenue Accounting Manager E55, Procurement Lead E57, Head of IT (Security) E56.
  • Hiring — brand/design/marketing. Motion Designer, Design Engineer, and Product Designer E31E32E33, Marketing Technology & Operations Lead (New York) E59, and EA to the Head of Marketing P12E10.
  • Hiring — geography. San Francisco is the hub; New York appears only for GTM roles E41E59.
  • Forks — thin, CI tooling only. The two confirmed forks are developer-CI actions: action-junit-report (fork of mikepenz/action-junit-report, Apache-2.0) P2 and run-report-action (fork of moonrepo/run-report-action, TypeScript) P3. A new (non-fork) repo terraform-provider-baseten appeared for infrastructure-as-code E54. There is no cited evidence of eval/model/agent framework forks in this pack.
  • Releases — agent tooling, CLI, and training surface. baseten-switch hit v0.5.0 → v0.5.1 (Beta) adding Claude Code native fallback controls, model picker, and attribution fixes P7P11E5E8; baseten-cli v0.4.0 shipped a full training surface (baseten train, baseten loops), model image build, audit logs, and autoscaling commands P18E17; truss shipped v0.18.28–v0.18.30rc0 (CUDA 12.9 default, JSON output, TRT-LLM LoRA cache settings, bnd config) P10P13P23E9E11E23; langchain-baseten libs/baseten v0.2.4 E34; run-report-action v1 P1.
  • Talking — performance frontier + infra engineering + open research. Posts claim leadership on Coval's voice AI benchmark (STT Pareto frontier, ~5× faster than OpenAI, Qwen3 ASR 1.7B Streaming) P5E3; explain delta weight syncs for managed RL rollouts (<40s, ~6s pause) P6E2; frame inference engineering via an "efficient frontier" P25E1; announce an MCP server + skill cutting agent task time/cost 7.5% on average P19E15; and launch Base Labs with open-publication research (continual learning, science of RL, BaseHub Data Foundry) W3W4. Enterprise changelog posts cover Viewer role, AWS AssumeRole, and Runtime OIDC P24P26E20E21E49.

Shipping

  • Model availability. GLM-5.3 and GLM-5.3-Flash via OpenAI-compatible Model APIs P27E27E35; DeepSeek V4 Pro 0813 (1.7T-param, MIT) W1E60; NVIDIA Nemotron 3.5 Lightning (30B MoE, 3B active) W2; Kimi K3, Whisper Large V3, and Qwen3.8-27B listed as popular models P24P26.
  • Developer tooling. baseten-switch Claude Code gateway P7P11; baseten-cli v0.4.0 P18; truss releases P10P13P23; MCP server + skill P19; terraform provider E54; langchain-baseten E34.
  • Enterprise/security features. Viewer role for read-only access P24, AWS AssumeRole for ECR/S3 pulls P26, Runtime OIDC E49, and CLI audit-log surfaces P18.

Research themes

  • Base Labs. Blue-sky research on continual learning and the science of RL, the BaseHub Data Foundry (open RL environments, training data, benchmarks), and "post-post training" deployed with a frontier safety stack W3W4.
  • Managed RL rollouts. Delta weight syncs for frontier open-weights models (GLM-5.3) across independent clusters P6.
  • Open-source post-training. Model selection and cost drivers (MoE active parameters, KV cache) across tiers P20.
  • Inference engineering. Efficient-frontier framing and DeepSeek V4 Pro 0813 harness/chat-template changes P25W1.

Hiring & scaling

  • Post-raise scaling. The $1.5B Series F (Altimeter, Conviction, Spark) is cited in nearly every job description P4P8P9P14P15P16P17P21P22P28.
  • Two simultaneous buildouts. (1) Infrastructure/capacity — Blackwell B200, multi-cloud orchestration, on-prem data centers P4P8P9; (2) commercialization/enterprise — IAM, finance systems, revenue ops, sales, marketing, and data analytics P15P16P17P21P22E29E30E41E42E48.
  • Training/RL org forming. AI Engineer, Forward Deployed Engineer (Training), and TPM Model Performance E25E47E50 align with the managed-rollout product P6.
  • Secondary hub. New York GTM roles E41E59.

Category implications

  • Capacity is becoming the neocloud moat. Hiring for fleet orchestration, multi-cloud workload movement, and Blackwell B200 deployment — and framing GPU fleet health as a unit-economics problem — implies neocloud competition is shifting from per-token pricing toward capital/capacity execution P4P8P9.
  • Inference → post-training/RL convergence. Managed rollouts, delta weight syncs, and baseten train/loops position Baseten to capture training-adjacent RL workloads, not just serving P6P18E47E50.
  • Enterprise trust surface. Fine-grained authorization for "agentic workloads," workload-based service-account credentials, Viewer role, AWS AssumeRole, and audit logs target regulated enterprise adoption (Harvey, HubSpot, Notion) P15P24P26P18.
  • Agent-native platform thesis. The MCP server plus Claude Code gateway (baseten-switch) treat coding agents as first-class platform operators, with a cited 7.5% average task cost/time reduction P19P7P11.
  • Open-source/research differentiation. Base Labs and open frontier-model availability (DeepSeek V4 Pro, Nemotron, GLM) suggest a strategy of capturing open-model demand and open-research mindshare W3W4W1W2P27.
  • Model supply speed as GTM wedge. Day-of/near-day availability of new frontier models (GLM-5.3, DeepSeek V4 Pro 0813, Nemotron 3.5 Lightning) is repeatedly surfaced in changelogs and blogs P27W1W2E60.

Traction highlights

  • $1.5B Series F led by Altimeter Capital, Conviction Partners, and Spark Capital P4.
  • Self-reported customer logos across posts: Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, Writer, Harvey, HubSpot, and Lovable P4P6P15.
  • Coval voice AI benchmark: STT ~5× faster than OpenAI with best WER (Qwen3 ASR 1.7B Streaming) P5.
  • Delta weight syncs under 40s with ~6s request pause P6; MCP 7.5% average task time/cost cut (up to ~57% for operation-heavy tasks) P19.
  • Frontier open-model availability: GLM-5.3/Flash, DeepSeek V4 Pro 0813, Nemotron 3.5 Lightning P27W1W2.
  • External discussion traction is thin — two blog posts show only 2 points/0 comments on HN E1E58, indicating these narratives are largely first-party.