- 标签:
- AI (232)
- Daily (205)
- Tech Trends (205)
- 周报 (30)
- Recommendation Systems (25)
- Weekly (25)
- Papers (25)
- 推荐系统 (16)
- 思考 (6)
- 论文 (6)
- Agentic Engineering (6)
- 日报 (5)
- 技术趋势 (5)
- 深度学习 (4)
- Harness Engineering (3)
- 推荐 (2)
- 工具 (2)
- 强化学习 (1)
- 思维模型 (1)
- Transformer (1)
- LLM (1)
- 管理 (1)
- 生成式 (1)
OpenAI made waves on multiple fronts: it terminated its Cursor partnership after SpaceX's acquisition, and reporters confirmed they've seen the next-gen Astra model. Meanwhile, GLM-5.3-Flash dominated the open-source conversation — Fireworks verified benchmark discrepancies before launch, and indepe
This week's narrative splits into two threads. The first is agent security moving from "theoretical risk" to "demonstrated attacks." OpenAI published its official postmortem of the HuggingFace intrusion, with critical takes from Gary Marcus and Zvi exposing problems that were less about sandbox hardness and more about missing monitoring and collective operational negligence. In the same week, Claude Code's default auto mode was broken — Johann Rehberger achieved roughly 80% attack success using a zip extraction plus malicious struct.py approach. Compounding this is the open-source supply chain: Anil Madhavapeddy reports an OCaml project faced exploit attempts within minutes of a patch discussion, and rclone received 40 security disclosures in one month — versus 20 over the previous decade. The second thread is open-weight models entering the "Day-0 inference engine support" era. On GLM-5.3's open-source release day, vLLM and SGLang shipped support simultaneously — SGLang even reused the runtime that generated its RL trajectories. Tencent's Hy4-preview likewise received vLLM day-0 support on release day. Unsloth compressed GLM-5.3 to 2-bit, shrinking 1.51TB to 239GB with roughly 81% precision retained. This means collaboration between open-source models and inference engines is now a default release-day action, not a community catch-up weeks later. Two major events in between deserve separate mention: NVIDIA acquiring HuggingFace for $13 billion, and OpenAI terminating model supply to Cursor following its acquisition by SpaceX. The former reshapes open-source model distribution; the latter marks the first time trust dynamics between model suppliers and downstream tools escalated into concrete contractual action.
The AI world is consolidating fast. NVIDIA reportedly moves to acquire Hugging Face for $12.9B — a seismic shift for open-source distribution. Meanwhile, GLM-5.3 and Tencent's Hy4 both dropped as massive open-weight MoE models, with Day-0 vLLM support. OpenAI cut off Cursor over SpaceX's acquisition
AI hit a security inflection point today: researchers broke Claude Code's auto mode via a zip-based attack, proving default safety settings aren't enough — sandboxing remains the only real defense. Meanwhile, Anthropic locked in a $45B compute deal with Nscale, and OpenAI joined 100+ organizations i
Open-source AI hit a new inflection point: Z.AI revealed the mysterious Ox Alpha as GLM-5.3-Flash — a 320B-A18B MoE with MIT license, trained at 1/9 the cost of Qwen3.7-Plus, and already proven on domestic Chinese AI chips with 3x inference efficiency gains. Alibaba countered with Qwen3.8-Flash (125
AI infrastructure hit a turning point today. NVIDIA unveiled the Vera Rubin NVL72 with 30x efficiency gains over GB300, while Meta countered with its own MTIA 300 training chip and MetaRoCE network protocol — the full-stack war is on. Hugging Face is exploring a $13B sale, nearly tripling its 2023 v
The "free lunch" era for AI agents is officially over. Anthropic's flagship Fable 5 model is struggling with adoption — just 8% share per Ramp's index — while the company's annualized revenue hits $65B. Drew Breunig's analysis crystallizes the shift: teams now route work by model tier, using GLM 5.2
AI hit a major milestone today: NVIDIA's AVO architecture scored a perfect 100% on ARC-AGI-3, proving system design — not just model scale — can unlock frontier-level performance. Anthropic is reportedly prepping a $100B+ IPO that could top SpaceX's record, with a $2T valuation target. Meta struck b
This week in AI, one clear theme dominates: the accelerating pace of capability gains is forcing evaluation, governance, and safety systems to speed up in tandem. Sam Altman announced a pause on some frontier RL training — the week's biggest industry event — citing capability advances outpacing the cadence of safety and alignment work. Meanwhile, NVIDIA's AVO agent completed all 183 tasks on ARC-AGI-3 with a 100% score, and Ornith-1.5 matched Claude Opus 4.8 through self-improvement training. Capability and safety are both accelerating — and pulling against each other. The second thread is a paradigm shift in evaluation and training environments: from static, zero-shot benchmarks toward real-time, long-horizon, self-generated environments. This week's FM-Bench (AnalogyAI) tests long-term decision-making via 20 years of football management simulation; EnvHarness (Google) makes static environments adaptively evolve; Wuying-Browser-Agent (Alibaba Cloud) introduces BrowserBench with an average of 37.9 steps. Evaluation is no longer asking "can it do it" — but "can it do it reliably over the long run." The third thread: inference optimization has entered a phase of fine-grained engineering. LFM2.5-DSpark (Liquid AI) delivers 3.2x speedup via speculative decoding; LMSYS's weight-caching daemon cuts engine loading from 495 seconds to 0.63 seconds. The cost curve is being pushed down on multiple fronts.
AI hit a major inflection point today: Z.ai CEO Tang Jie declared "parameter count is dead," crediting GLM 5.3's leap entirely to long-horizon RL training in synthetic environments. NVIDIA dropped $6B on Poolside's "model factory," while Moderna/Merck's personalized mRNA cancer vaccine hit Phase III