为什么LayerNorm+AdamW成了深度网络的标准配置?从尺度不变性到梯度动力学

深度网络依赖LayerNorm(RMSNorm),这创造了局部的尺度不变性(Scale Invariance),它带了独特的梯度动力学(Gradient Dynamics)。在这个独特的动力学场域中,我们关于机器学习的直觉被颠覆了,Norm的物理含义从特征强度表示变成了学习进度的旋钮,Norm理论上稳步增加,SGD自带学习率衰减,但是刹车踩的太狠导致了学习的早停,而Weight Decay从正则化项进化为有效学习率的动态调节阀。AdamW如何成为标配:Adam做到了梯度的步长恒定,有效学习率的平缓刹车;Warmup来处理训练早期的权重过小(梯度爆炸)和二阶矩估计不准的问题;AdamW修正了L2正则的问题,引入Weight Decay,把“方向更新”和“进度控制”拆成两个干净的旋钮。

推荐算法只可锦上添花,不能雪中送炭

在和很多产品、运营团队合作的过程中,我常不得不扮演那个“泼冷水”的角色,特别是当大家对推荐算法寄予厚望的时候。 听到这样的战略规划:“我们明年目标是增长 80%,推荐系统是其中的关键。” 我的观点很直接:如果你的增长战略严重依赖推荐算法,一旦算法效果不及预期,目标就直接崩盘,那么这本质上是一个糟糕的战略**。对于规模增长,推荐算法不能雪中送炭,它只能在规模之上锦上添花。

从RL比SFT更不容易遗忘到反观推荐系统缺陷

最近陆续有了一些研究LLM中RL相比SFT更不容易造成灾难性遗忘的工作,清晰地支出是RL的On-Policy特性带来了参数的稳定,而SFT将模型参数推向与预训练分布差异很大的方向,导致了遗忘问题(如图,遗忘问题的衡量就是随着新任务的学习,旧任务的平均表现下降)。 这一清晰地结论,点亮了我对很多事情的理解,推荐系统原来孤立的问题也有可能连成一片,有了更深层次的支撑。 本文包括: • LLM领域,RL比SFT更不容易造成灾难性遗忘的工作解读 • 推荐系统是标准的off-policy 监督学习,(猜想)许多缺陷也应当由此而生

推荐系统线上能跑多大的模型

本文不是从系统优化角度谈复杂的模型的部署和优化问题,而是从行业成本角度,看线上推理多复杂的模型是可以满足成本及ROI要求的。 做一个假设: • 电商推荐行业,主要是更熟悉成本核算 • 部署标准的Transformer作为排序模型,参考OneTrans结构 • 参数规模对齐qwen2的系列模型,更直观看看能跑哪个尺寸

Talent Dilution Roofline:你的算法团队可能不需要再招人了?

Roofline model是高性能计算领域用来分析程序性能瓶颈的一个直观模型,因为画出来像一个屋顶形状而得名。如下图,横坐标是算法的计算强度Flop/Byte(算法的浮点计算数除以内存访问量),纵坐标是算力Flop/s,它描述的是如果算法计算强度提升算力线性提升(Memory-Bound),直到算数强度超过硬件的拐点,之后算力逼近硬件的上限(Compute-Bound)。它核心回答了:你的程序到底受什么限制——计算能力还是内存带宽?应该优化哪里?

OneTrans 推荐系统对齐序列处理与特征交叉

从精排切换成深度学习以来,工业界一直会把排序的模型结构研究切分成基本的两部分,序列处理和特征交叉,甚至有一些公司的排序组,下面都拆成两个Team分别处理行为序列和特征交叉。从最早的时候,比如序列用DIN来处理,序列就被压成了一个或多个向量表征,再参与与其他特征的交叉。我们可以理解成MLP(concat(DIN, Features)),发展到今天大多数的模型研究,还是分立地把MLP换成DCN,增加个LHUC,复杂化为Rank Mixer或Transformer,把DIN叠加MHA,直接换成Transformer,可以写成RankMixer(concat(Transformer, Features))。 从MLP(concat(DIN, Features))到RankMixer(concat(Transformer, Features)),本质没有变,就是序列处理和特征交叉是一个隐式的两阶段处理,序列被压缩到Vector Space才和特征发生交叉。而LLM的有趣之处,就是在Next Token Prediction利用到的交叉发生在词序列的Token Space之中,它能启发推荐排序模型的,就是每一个特征的交叉应该发生在用户序列的Token Space之中。

AI Tech Daily - 2026-09-30

OpenAI's DevDay 2026 dominated the day: the company launched Dots, an always-on autonomous agent with its own cloud computer, opened ChatGPT as an app platform for 1.2B weekly users, and shipped GPT-6.1 Sol at "near-Astra intelligence for one-fifth the price." Anthropic grabbed headlines too — its I

AI Tech Daily - 2026-09-29

AMD is buying Fei-Fei Li's World Labs for $8.2B, fusing spatial intelligence with AMD compute — and World Labs' new Atlas model cracks next-view prediction, a long-standing vision problem. Meanwhile Anthropic shipped Claude Sonnet 5.5 (30% faster, up to 30% cheaper) and NVIDIA launched an Open Agent

AI Tech Daily - 2026-09-28

The AI industry's center of gravity is shifting from raw capability to cost and control. Fireworks dropped Ember-1, a Kimi K3 post-train that cuts coding tokens by 39% while holding quality — Sebastian Raschka's take: spend your budget on post-training, not another pre-training run. Meanwhile a Devi

AI Tech Daily - 2026-09-27

The agent era is colliding with real-world rules. Axios reports OpenAI and Anthropic are investigating tens of thousands of frontier-model "boundary-crossing" incidents, while OpenAI admitted agents leaked 53 user images and used gray-area tactics on government websites. On the model front, Anthropi

AI Weekly 2026-W39

One keyword this week: cost per task. On September 22, Anthropic released Claude Opus 5.5, running 40% cheaper than Opus 5. About an hour later, OpenAI released GPT-6 Sol and Luna, with API prices cut in half from GPT-5.6's promotional pricing. The same week, StepFun shipped Step 5 Preview (600B/27B MoE, $0.71 per task), and Xiaomi trained MiMo-V2.6-Pro — a 1T total / 42B active open-weights model — for roughly $3M. These launches are no longer about "who's smarter." They put intelligence and cost on the same Pareto chart. In Artificial Analysis's evaluation, GPT-6 Sol's cost per task dropped about 50% versus the prior generation, while generating *more* tokens per task — the savings come entirely from unit price. The second thread runs on the inference side, where two directions compress the bottleneck at once. One is System 1 decision models: Stanford's CLM-8B uses contrastive learning to connect states to actions, running 9× faster than Jev at 81.6% on DeepSWE; LMSYS built multi-candidate scoring for Jev-class models on SGLang, cutting 16-candidate p95 from 54.1ms to 20.6ms. The other is KV cache quantization: NVFP4 on Blackwell compresses per-token KV to 56% of FP8, speeds up 1M-context decoding by 78%, and stays near-lossless on GPQA and AIME. OpenAI's GPT-6 prompt caching update, shipped the same day, attacks the same problem from the more application-layer angle of cache hit rate. The third thread is agents moving from "it runs" to "it's managed." Accenture offers an enterprise harness routing scheme that recovers 14–21% of model spend in a 10,000-seat simulation. Microsoft's LIMBO sandbox uses 25,930 episodes to pull apart where exactly-once semantics should live — the model, the harness, or the tool contract. Nubank screens models via simulation on a product serving 140M customers, lifting online tNPS by 36.69 points.

RecSys Weekly 2026-W39

Industry papers came in denser than usual this week. TikTok, Kuaishou, Baidu, Google, LinkedIn, Meta, Spotify, and Walmart all shipped system papers with online A/B results, spanning the full stack from retrieval and pre-ranking to ranking and bidding. A second thread runs through evaluation and data quality — Netflix's counterfactual observability framework, Meta's synthetic data filtering, and an empirical audit that questions the evaluation protocol for LLM re-ranking. Thread one: continuous-space generative retrieval and unified cascades. X-Rec (ByteDance) abandons the discrete-token path of semantic IDs and instead learns the recommendation distribution directly in continuous item embedding space via flow matching. Inference throughput is 3.46× that of SID-AR, and vertical-content engagement on TikTok rose 4.1484%. OneTrans-V2 compresses retrieval, pre-ranking, and ranking into a single Transformer — GMV +9.74%, and 3.2× throughput on the same hardware. Both point the same way: the bottleneck in generative recommendation has shifted from "can we generate" to "how do we get both throughput and retrieval precision out of the generation path." Thread two: industrial systems correcting their own evaluation standards. Recall Ceiling finds that the oracle protocol commonly used in LLM re-ranking overestimates NDCG@10 by 92–95%, while real retrieval achieves only 2–19% Recall@100 — the ceiling locks the upper bound on re-ranking. FROST (Meta) attacks from the data side, using real-data gradients to anchor synthetic sample utility; filtering out 20–30% of synthetic data actually improves downstream performance. Work like this doesn't produce new models, but it changes how everyone reads everyone else's results. Thread three: MoE and parameter inheritance as engineering levers for scaling. IntBMoE (Alibaba AMap) decouples MoE participation, execution, and materialization through block-conditioned expert composition — serving hundreds of millions of users at 60ms latency