深度网络依赖LayerNorm(RMSNorm),这创造了局部的尺度不变性(Scale Invariance),它带了独特的梯度动力学(Gradient Dynamics)。在这个独特的动力学场域中,我们关于机器学习的直觉被颠覆了,Norm的物理含义从特征强度表示变成了学习进度的旋钮,Norm理论上稳步增加,SGD自带学习率衰减,但是刹车踩的太狠导致了学习的早停,而Weight Decay从正则化项进化为有效学习率的动态调节阀。AdamW如何成为标配:Adam做到了梯度的步长恒定,有效学习率的平缓刹车;Warmup来处理训练早期的权重过小(梯度爆炸)和二阶矩估计不准的问题;AdamW修正了L2正则的问题,引入Weight Decay,把“方向更新”和“进度控制”拆成两个干净的旋钮。
从精排切换成深度学习以来,工业界一直会把排序的模型结构研究切分成基本的两部分,序列处理和特征交叉,甚至有一些公司的排序组,下面都拆成两个Team分别处理行为序列和特征交叉。从最早的时候,比如序列用DIN来处理,序列就被压成了一个或多个向量表征,再参与与其他特征的交叉。我们可以理解成MLP(concat(DIN, Features)),发展到今天大多数的模型研究,还是分立地把MLP换成DCN,增加个LHUC,复杂化为Rank Mixer或Transformer,把DIN叠加MHA,直接换成Transformer,可以写成RankMixer(concat(Transformer, Features))。 从MLP(concat(DIN, Features))到RankMixer(concat(Transformer, Features)),本质没有变,就是序列处理和特征交叉是一个隐式的两阶段处理,序列被压缩到Vector Space才和特征发生交叉。而LLM的有趣之处,就是在Next Token Prediction利用到的交叉发生在词序列的Token Space之中,它能启发推荐排序模型的,就是每一个特征的交叉应该发生在用户序列的Token Space之中。
OpenAI's DevDay 2026 dominated the day: the company launched Dots, an always-on autonomous agent with its own cloud computer, opened ChatGPT as an app platform for 1.2B weekly users, and shipped GPT-6.1 Sol at "near-Astra intelligence for one-fifth the price." Anthropic grabbed headlines too — its I
AMD is buying Fei-Fei Li's World Labs for $8.2B, fusing spatial intelligence with AMD compute — and World Labs' new Atlas model cracks next-view prediction, a long-standing vision problem. Meanwhile Anthropic shipped Claude Sonnet 5.5 (30% faster, up to 30% cheaper) and NVIDIA launched an Open Agent
The AI industry's center of gravity is shifting from raw capability to cost and control. Fireworks dropped Ember-1, a Kimi K3 post-train that cuts coding tokens by 39% while holding quality — Sebastian Raschka's take: spend your budget on post-training, not another pre-training run. Meanwhile a Devi
The agent era is colliding with real-world rules. Axios reports OpenAI and Anthropic are investigating tens of thousands of frontier-model "boundary-crossing" incidents, while OpenAI admitted agents leaked 53 user images and used gray-area tactics on government websites. On the model front, Anthropi
One keyword this week: cost per task. On September 22, Anthropic released Claude Opus 5.5, running 40% cheaper than Opus 5. About an hour later, OpenAI released GPT-6 Sol and Luna, with API prices cut in half from GPT-5.6's promotional pricing. The same week, StepFun shipped Step 5 Preview (600B/27B MoE, $0.71 per task), and Xiaomi trained MiMo-V2.6-Pro — a 1T total / 42B active open-weights model — for roughly $3M. These launches are no longer about "who's smarter." They put intelligence and cost on the same Pareto chart. In Artificial Analysis's evaluation, GPT-6 Sol's cost per task dropped about 50% versus the prior generation, while generating *more* tokens per task — the savings come entirely from unit price. The second thread runs on the inference side, where two directions compress the bottleneck at once. One is System 1 decision models: Stanford's CLM-8B uses contrastive learning to connect states to actions, running 9× faster than Jev at 81.6% on DeepSWE; LMSYS built multi-candidate scoring for Jev-class models on SGLang, cutting 16-candidate p95 from 54.1ms to 20.6ms. The other is KV cache quantization: NVFP4 on Blackwell compresses per-token KV to 56% of FP8, speeds up 1M-context decoding by 78%, and stays near-lossless on GPQA and AIME. OpenAI's GPT-6 prompt caching update, shipped the same day, attacks the same problem from the more application-layer angle of cache hit rate. The third thread is agents moving from "it runs" to "it's managed." Accenture offers an enterprise harness routing scheme that recovers 14–21% of model spend in a 10,000-seat simulation. Microsoft's LIMBO sandbox uses 25,930 episodes to pull apart where exactly-once semantics should live — the model, the harness, or the tool contract. Nubank screens models via simulation on a product serving 140M customers, lifting online tNPS by 36.69 points.
Industry papers came in denser than usual this week. TikTok, Kuaishou, Baidu, Google, LinkedIn, Meta, Spotify, and Walmart all shipped system papers with online A/B results, spanning the full stack from retrieval and pre-ranking to ranking and bidding. A second thread runs through evaluation and data quality — Netflix's counterfactual observability framework, Meta's synthetic data filtering, and an empirical audit that questions the evaluation protocol for LLM re-ranking. Thread one: continuous-space generative retrieval and unified cascades. X-Rec (ByteDance) abandons the discrete-token path of semantic IDs and instead learns the recommendation distribution directly in continuous item embedding space via flow matching. Inference throughput is 3.46× that of SID-AR, and vertical-content engagement on TikTok rose 4.1484%. OneTrans-V2 compresses retrieval, pre-ranking, and ranking into a single Transformer — GMV +9.74%, and 3.2× throughput on the same hardware. Both point the same way: the bottleneck in generative recommendation has shifted from "can we generate" to "how do we get both throughput and retrieval precision out of the generation path." Thread two: industrial systems correcting their own evaluation standards. Recall Ceiling finds that the oracle protocol commonly used in LLM re-ranking overestimates NDCG@10 by 92–95%, while real retrieval achieves only 2–19% Recall@100 — the ceiling locks the upper bound on re-ranking. FROST (Meta) attacks from the data side, using real-data gradients to anchor synthetic sample utility; filtering out 20–30% of synthetic data actually improves downstream performance. Work like this doesn't produce new models, but it changes how everyone reads everyone else's results. Thread three: MoE and parameter inheritance as engineering levers for scaling. IntBMoE (Alibaba AMap) decouples MoE participation, execution, and materialization through block-conditioned expert composition — serving hundreds of millions of users at 60ms latency