深度网络依赖LayerNorm(RMSNorm),这创造了局部的尺度不变性(Scale Invariance),它带了独特的梯度动力学(Gradient Dynamics)。在这个独特的动力学场域中,我们关于机器学习的直觉被颠覆了,Norm的物理含义从特征强度表示变成了学习进度的旋钮,Norm理论上稳步增加,SGD自带学习率衰减,但是刹车踩的太狠导致了学习的早停,而Weight Decay从正则化项进化为有效学习率的动态调节阀。AdamW如何成为标配:Adam做到了梯度的步长恒定,有效学习率的平缓刹车;Warmup来处理训练早期的权重过小(梯度爆炸)和二阶矩估计不准的问题;AdamW修正了L2正则的问题,引入Weight Decay,把“方向更新”和“进度控制”拆成两个干净的旋钮。
从精排切换成深度学习以来,工业界一直会把排序的模型结构研究切分成基本的两部分,序列处理和特征交叉,甚至有一些公司的排序组,下面都拆成两个Team分别处理行为序列和特征交叉。从最早的时候,比如序列用DIN来处理,序列就被压成了一个或多个向量表征,再参与与其他特征的交叉。我们可以理解成MLP(concat(DIN, Features)),发展到今天大多数的模型研究,还是分立地把MLP换成DCN,增加个LHUC,复杂化为Rank Mixer或Transformer,把DIN叠加MHA,直接换成Transformer,可以写成RankMixer(concat(Transformer, Features))。 从MLP(concat(DIN, Features))到RankMixer(concat(Transformer, Features)),本质没有变,就是序列处理和特征交叉是一个隐式的两阶段处理,序列被压缩到Vector Space才和特征发生交叉。而LLM的有趣之处,就是在Next Token Prediction利用到的交叉发生在词序列的Token Space之中,它能启发推荐排序模型的,就是每一个特征的交叉应该发生在用户序列的Token Space之中。
The AI price war just escalated again. Google launched Gemini 3.7 Flash at half the token cost with big benchmark jumps, while OpenAI previewed Ultrafast — a Cerebras-powered tier that runs GPT-5.6 Sol 14x faster at up to 750 tokens/sec. Meanwhile, xAI's Grok 4.6 hit Perplexity at 60% lower cost, an
Frontier Model Day reshaped the competitive landscape: xAI shipped Grok 4.6 (1.5T params) at $2/$6 per million tokens — roughly 60% cheaper than Claude Opus 5 — while Alibaba open-sourced Qwen3.8-Max (2.4T total, 95B active) with day-0 vLLM support. DeepSeek countered with V4-Pro 0813, topping Termi
AI hit a commercial inflection point today. OpenAI began testing ads in ChatGPT across six markets, while Anthropic canceled a planned price hike — the subscription-only era is ending. Meanwhile, River AI raised $1.1B to build "personally owned AI," and Gemini crossed 1B monthly users, making it Goo
AI hit a major infrastructure milestone today: NVIDIA teamed up with six Wall Street giants — Apollo, Blackstone, BlackRock, Brookfield, Goldman Sachs, and KKR — to build a $500B+ compute financing platform, turning AI chips into a new asset class. Meta open-sourced Muse Glimmer 30B under Apache 2.0
The AI safety debate hit a new peak today: CNBC revealed that OpenAI, Anthropic, and Meta's recent model "runaway" incidents all trace back to the same Israeli startup, Irregular — a red-team testing vendor backed by Sequoia and Redpoint. Meanwhile, Australia saw its first autonomous AI attack, with