Neolabfresh 3w

Meituan (LongCat)

Signal timeline56 total

Nothing in this view yet.

Top signals

  1. #1Modelsmeituan-longcat/LongCat-Flash-Chat7.0
  2. #2Modelsmeituan-longcat/LongCat-Flash-Thinking-26017.0
  3. #3Modelsmeituan-longcat/LongCat-Image-Edit7.0
  4. #4Reposmeituan-longcat/LongCat-2.06.0
  5. #5Reposmeituan-longcat/LongCat-AudioDiT6.0

Agent answer

Meituan (LongCat) has 56 loaded public signals: 0 hiring, 6 forks, 19 releases or model cards, 0 talking, and 31 repos. Latest signal: meituan-longcat/LongCat-2.0. Data-business radar is currently scoped to frontier labs, so this category does not expose radar lanes. The standing analysis was generated with deepseek-v4-pro and 89 evidence refs.

Meituan (LongCat)

has loaded 56 public signals

Meituan (LongCat)

has hiring signal count 0

Meituan (LongCat)

has fork signal count 6

Meituan (LongCat)

has release signal count 19

Analysis — agent synthesisfull report →generated July 4, 2026

Thesis

Meituan's LongCat lab is executing a full-spectrum, model-system co-design strategy that spans text, image, video, audio, omni-modal generation, formal reasoning, and agentic evaluation — all anchored by a growing family of open-weight Mixture-of-Experts models. The capstone is LongCat-2.0, a 1.6T-parameter MoE with ~48B activated parameters trained and deployed end-to-end on domestic Chinese AI ASIC superpods, which the lab claims is the first trillion-plus-parameter model to achieve this milestone P3W2W3. This domestic silicon bet — framed explicitly as a response to US export controls W4 — is paired with deep infrastructure work: custom inference engines (SGLang-FluentLLM), distributed KV cache frameworks (Omni-Flow), and kernel-level optimizations forked from DeepSeek and FlashInfer P1P2P22E50E54E55E56. The lab simultaneously builds evaluation moats through an array of benchmarks (AMO-Bench, VitaBench, R-HORIZON, UNO-Bench, Meeseeks, MineExplorer, WBench, General365, LARYBench, SOP-Maze, PIEbench) that have gained traction with external model teams including Qwen and ByteDance Seed P11E37. The intensity and breadth of output — spanning basic research through production inference tooling — suggests a well-resourced organization pursuing a vertically integrated AI capability from silicon to application.

Signal desks

Hiring

  • Meituan unveiled a LongCat AI Model Internship Program aimed at cultivating AGI talent, though the evidence pack lacks detailed role-level job listings, team breakdowns, or location-specific hiring signals W5. No cited open roles, team expansions, or location hubs beyond this single announcement.

Forks

  • meituan-longcat/FlashMLA — Fork of deepseek-ai/FlashMLA, with custom swapAB and FP8 KVCache + FP8 Compute optimizations applied E50P22. Language: C++/CUDA. Technical theme: MLA attention kernel optimization for long-context decoding.
  • meituan-longcat/DeepGEMM — Fork of deepseek-ai/DeepGEMM, with SwapAB Offset + PDL optimizations E54P22. Language: C++/CUDA. Technical theme: GEMM kernel optimization for MoE inference.
  • meituan-longcat/DeepEP — Fork of deepseek-ai/DeepEP E55. Language: C++/CUDA. Technical theme: expert parallelism communication for MoE models.
  • meituan-longcat/flashinfer — Fork of flashinfer-ai/flashinfer, with communication–computation fused kernel optimizations on a feature/longcat_main branch E56P22. Language: Python/C++/CUDA. Technical theme: inference kernel fusion for attention and sampling.
  • meituan-longcat/fast-hadamard-transform — Fork of Dao-AILab/fast-hadamard-transform E53. Language: Python/C++/CUDA. Technical theme: quantization-friendly linear algebra primitives.
  • meituan-longcat/mscclpp — Fork of microsoft/mscclpp E51. Language: C++/CUDA. Technical theme: multi-node collective communication for distributed training/inference.

Releases

  • LongCat-Flash-Chat — 560B-parameter MoE chat model released Aug 2025 on HuggingFace (80,386 downloads, 536 likes) and GitHub (1,339 stars) E1E23P8. MIT license. Text-generation pipeline. Companion arxiv tech report P8.
  • LongCat-Flash-Thinking — 560B-parameter large reasoning model released Sep 2025 (149 likes, 71 downloads on HF; 285 GitHub stars) E8E28P7. MIT license. Text-generation pipeline.
  • LongCat-Flash-Omni — ~560B-parameter any-to-any omni-modal model released Oct 2025 (114 likes, 31 downloads; 492 GitHub stars) E10E32P14. MIT license.
  • LongCat-Video — Text-to-video generation model released Oct 2025 (531 likes, 2,328 downloads; 4,273 GitHub stars, 669 forks) E2E26P16. MIT license. Includes LongCat-Video-Avatar variant.
  • LongCat-Audio-Codec — Audio tokenizer/detokenizer for speech LLMs released Oct 2025 (43 likes; 301 GitHub stars) E16E35P9. MIT license.
  • LongCat-Image family — Text-to-image (LongCat-Image: 17,669 downloads, 247 likes), image editing (LongCat-Image-Edit: 23,724 downloads, 183 likes), dev variant, and turbo variant released Dec 2025–Feb 2026 E3E7E12E13P18. Apache-2.0.
  • LongCat-Flash-Thinking-2601 — Updated reasoning model released Jan 2026 (115 likes, 2,721 downloads; 254 GitHub stars) E9E36P20. MIT license.
  • LongCat-Flash-Lite — Smaller 69B-parameter model released Jan 2026 (192 likes, 1,625 downloads) E6. MIT license.
  • LongCat-Flash-Thinking-ZigZag — Reasoning variant released Jan 2026 (32 likes; 8 GitHub stars) E21E46P21. MIT license.
  • LongCat-HeavyMode-Summary — 560B-parameter summary variant released Jan 2026 (13 likes) E27. MIT license.
  • LongCat-Flash-Prover — 560B formal theorem proving model (Lean4) released Mar 2026 (34 likes, 157 downloads; 85 GitHub stars) E19E39P25. MIT license.
  • LongCat-Next — Native multimodal model (~74B params, any-to-any pipeline) released Mar 2026 (193 likes, 619 downloads; 438 GitHub stars) E5E33P26. MIT license.
  • LongCat-AudioDiT — Diffusion TTS models (1B and 3.5B) released Mar 2026 (75 likes for 3.5B, 46 likes for 1B; 522 GitHub stars) E11E14E30P28. MIT license.
  • LongCat-2.0 — 1.6T MoE model released Jun 2026 (167 likes; 41 GitHub stars at time of snapshot) E4E15P3P4. MIT license.
  • Benchmark releases: VitaBench (ICLR 2026, 145 stars) P11E37; AMO-Bench (128 stars) P15E38; UNO-Bench (78 stars) P13E41; R-HORIZON (ICLR 2026, 26 stars) P12E44; Meeseeks (57 stars) P6E42; MineExplorer (11 stars) P5E22; WBench (159 stars) E25; VitaBench-2.0 (52 stars) E24; General365 (81 stars) E34; LARYBench (151 stars) E31; SOP-Maze (7 stars) P10E45; PIEbench (4 stars) P19E48; DPT-Agent (ACL 2025, 5 stars) P17E47.
  • Infrastructure releases: omni-flow — multimodal inference orchestration framework with distributed KV cache (7 stars, Jun 2026) P1E17; omni-flow-sglang — SGLang modifications for Omni-Flow compute module (2 stars, Jun 2026) P2E18; SGLang-FluentLLM — custom inference engine with speculative decoding, CUDA graph fusion, layer-wise KVCache transfer, and Decode Radix Tree Cache (83 stars, Feb 2026) P22E40; LongCat-Next-inference — multimodal inference testing framework (31 stars, Mar 2026) P27E43; triton_configs — Triton inference server configuration (Feb 2026) P23E52; eps — C++ repo with no published README (2 stars, Feb 2026) P24E49.

Talking

  • Domestic chip narrative dominates press: VentureBeat coverage frames LongCat-2.0 as "near-frontier agentic coding model ... trained entirely on Chinese chips" that has been leading OpenRouter W1. CNA reports Meituan claims performance comparable to Google's Gemini 3.1 Pro W2. SiliconANGLE highlights the ASIC superpod training origin and reduced dependence on Nvidia W3. TNW positions the release as "a pointed answer to US export controls" W4.
  • Talent cultivation: xix.ai reports on a LongCat AI Model Internship Program for AGI talent W5.
  • No cited evidence of direct lab blog posts, social media threads, or community discussion content beyond repository READMEs and press coverage.

Shipping

LongCat has maintained a sustained release cadence spanning approximately 10 months (Aug 2025–Jun 2026), shipping at least 15 distinct model artifacts across HuggingFace and numerous supporting repositories. The release tempo accelerated in late 2025 (Flash-Chat, Flash-Thinking, Flash-Omni, Video, Audio-Codec, Image all within Aug–Dec 2025), continued in Q1 2026 with iterative reasoning model updates (2601, ZigZag, Prover, Lite, HeavyMode), and pivoted toward native multimodal and audio generation in late Q1 (LongCat-Next, AudioDiT). The capstone LongCat-2.0 landed at end of Q2 2026 E4P3.

Infrastructure shipping runs in parallel: SGLang-FluentLLM (Feb 2026) exposed speculative decoding with Eagle/MTP/PLD support, layer-wise KVCache transfer, and Decode Radix Tree Cache — indicating production-scale inference deployment needs P22. Omni-Flow (Jun 2026) shipped a three-layer orchestration framework for multimodal inference with distributed KV cache sharing, targeting the heterogeneous compute challenge of serving diffusion models alongside LLMs on the same SGLang-based forward path P1P2. LongCat-Next-inference (Mar 2026) shipped a unified Hidden State interface for native multimodal models with a decoupled generation pipeline P27.

Benchmark shipping is also extensive: 12+ evaluation repos released, with VitaBench P11 and R-HORIZON P12 both accepted at ICLR 2026, DPT-Agent at ACL 2025 P17, and VitaBench cited by Qwen3.5 and ByteDance Seed2.0 P11.

Research themes

1. Extreme-scale MoE with domestic silicon: LongCat-2.0 represents a 1.6T total / ~48B activated MoE trained entirely on AI ASIC superpods of 50,000 domestically made chips, demonstrating training-at-scale capability outside Nvidia ecosystems P3W2W3. This builds on the earlier 560B MoE architecture introduced with LongCat-Flash-Chat P8.

2. Model-system co-design for inference: The lab explicitly states this as a guiding principle P22. Research outputs include speculative decoding workflow refactoring for overlap scheduling, CUDA graph fusion of target/verify/draft, layer-wise KVCache transfer across PD boundaries, Decode Radix Tree Cache for KVCache transfer volume reduction P22, and the three-layer Omni-Flow architecture for multimodal inference orchestration with distributed KV cache P1.

3. Native multimodal architectures: LongCat-Next implements a unified Hidden State interface where all modality-specific inputs are converted to a uniform representation before entering the LLM core, enabling a single forward path for image understanding, image generation, audio-to-text, speech synthesis, and audio-to-audio tasks P27. LongCat-Flash-Omni extends this to any-to-any generation P14.

4. Formal reasoning and theorem proving: LongCat-Flash-Prover targets native formal reasoning in Lean4 for mathematics formalization and proving via agentic tool-integrated reasoning P25. R-HORIZON studies long-horizon reasoning degradation in LRMs through query composition, accepted at ICLR 2026 P12.

5. Diffusion-based generation across modalities: LongCat-Video for text-to-video P16, LongCat-Image for text-to-image with editing capabilities P18, and LongCat-AudioDiT for high-fidelity TTS in waveform latent space P28, all published with arxiv tech reports.

6. Comprehensive evaluation science: The lab produces benchmarks spanning math (AMO-Bench) P15, omni-modal composition (UNO-Bench) P13, agentic tool-use (VitaBench, DPT-Agent) P11P17, instruction following (Meeseeks) P6, open-world agent exploration (MineExplorer) P5, business SOP execution (SOP-Maze) P10, video world models (WBench) E25, general reasoning (General365) E34, and multilingual QA (PIEbench) P19.

7. Kernel-level optimization: SnapMLA for FP8-quantized MLA decoding P22, FlashMLA swapAB and FP8 KVCache optimizations E50P22, DeepGEMM swapAB offset optimizations E54P22, and communication-computation fused kernels in FlashInfer E56P22.

Hiring & scaling

Evidence is thin in this pack. Meituan announced a LongCat AI Model Internship Program to "cultivate AGI talent" W5, but no cited evidence provides granular detail on specific roles, team composition, headcount, location strategy, or organizational structure. The rate and breadth of technical output (15+ models, 12+ benchmarks, 6+ infrastructure repos across Aug 2025–Jun 2026) implies a substantial research organization, but the pack lacks direct hiring data to confirm scale. No evidence on data labeling workforce, evaluation contractors, or commercialization hiring.

Category implications

Infrastructure and silicon strategy: The claim that LongCat-2.0 is the first trillion-plus-parameter model fully trained and deployed on domestic Chinese AI ASIC superpods P3W2W3 has significant supply-chain implications. If verified through open-source community benchmarking W4, it demonstrates that large-scale frontier training is viable without Nvidia GPUs, potentially reshaping hardware procurement strategies for Chinese AI labs and reducing the effective impact of US export controls W4. The ASIC superpod architecture referenced in the model card P3 indicates custom silicon deployment at scale, not merely consumer-grade alternatives.

Inference infrastructure as competitive moat: The SGLang-FluentLLM and Omni-Flow releases suggest LongCat is investing in proprietary inference infrastructure that goes beyond model weights. The combination of speculative decoding with overlap scheduling, CUDA graph fusion, layer-wise KVCache transfer, and a three-tier paged KV cache hierarchy (GPU→CPU→SSD) P22P1 positions them to serve multimodal workloads at production latency. This blurs the line between model provider and inference platform, with implications for how other labs using SGLang P2 might adopt or compete with these optimizations.

Multimodal product surface: The breadth of modality coverage — text, image generation, image editing, video generation, speech TTS, audio codec, omni-modal — paired with native multimodal architectures like LongCat-Next P27P26, suggests a product strategy targeting content creation, media, and interactive applications. The LongCat-Video-Avatar variant P16 additionally points toward digital human/persona applications.

Agent and evaluation ecosystem: The lab's heavy investment in benchmarks — especially those gaining external adoption like VitaBench (cited by Qwen and ByteDance) P11 and AMO-Bench (used to track Gemini, Qwen, Kimi-K2, GLM progress) P15 — positions LongCat as an evaluation standard-setter. This creates network effects where external labs optimize against LongCat-designed metrics. MineExplorer P5 and DPT-Agent P17 extend this into embodied and collaborative agent evaluation.

Open-weight licensing strategy: Consistent MIT licensing across nearly all model releases P3P7P8P14P16P20P21P25P26P28 (with Apache-2.0 for Image models P18) suggests a permissive open-source posture that maximizes adoption and community contribution while the parent company (Meituan) likely monetizes through its core delivery and services platform rather than model API fees.

Formal reasoning differentiation: LongCat-Flash-Prover targeting Lean4 formal mathematics P25 is a relatively rare focus among large labs and could signal intent to build verifiable reasoning capabilities relevant to code verification, mathematical research, or safety-critical AI applications.

Traction highlights

  • GitHub stars: LongCat-Video leads at 4,273 stars with 669 forks P16; LongCat-Flash-Chat at 1,339 stars P8; LongCat-Image at 695 stars P18; LongCat-AudioDiT at 522 stars P28; LongCat-Flash-Omni at 492 stars P14; LongCat-Next at 438 stars P26; LongCat-Audio-Codec at 301 stars P9; LongCat-Flash-Thinking at 285 stars P7; LongCat-Flash-Thinking-2601 at 254 stars P20. Multiple benchmark repos exceed 100 stars: WBench (159), LARYBench (151), VitaBench (145), AMO-Bench (128), General365 (81), UNO-Bench (78) E25E31E37E38E34E41.
  • HuggingFace downloads: LongCat-Flash-Chat at 80,386 E1; LongCat-Image-Edit at 23,724 E7; LongCat-Image-Edit-Turbo at 22,419 E12; LongCat-Image at 17,669 E3; LongCat-AudioDiT-3.5B at 5,909 E11; LongCat-Flash-Thinking-2601 at 2,721 E9; LongCat-Video at 2,328 E2.
  • Academic recognition: Two ICLR 2026 acceptances (VitaBench P11, R-HORIZON P12), one ACL 2025 acceptance (DPT-Agent P17).
  • External benchmark adoption: VitaBench cited by Qwen3.5, Seed2.0, and Qwen3-Max-Thinking evaluations P11. AMO-Bench leaderboard tracks SOTA from Gemini 3 Pro, Kimi-K2-Thinking, Qwen3-Max-Thinking, and GLM-4.7 P15.
  • Press coverage: VentureBeat W1, CNA W2, SiliconANGLE W3, and TNW W4 all covered the LongCat-2.0 domestic-chip narrative within 24 hours of release.
  • OpenRouter presence: LongCat-2.0 reported as leading OpenRouter for agentic coding, per VentureBeat W1.
  • HuggingFace likes: LongCat-Flash-Chat at 536 E1; LongCat-Video at 531 E2; LongCat-Image at 247 E3; LongCat-Next at 193 E5; LongCat-Flash-Lite at 192 E6; LongCat-Image-Edit at 183 E7; LongCat-2.0 at 167 E4; LongCat-Flash-Thinking at 149 E8.