这一期,我们来当一回AI的“监考官”和“心理医生”,看看怎么科学地判断AI是在“背课文”还是“会造句”。我们还会探究,当AI说自己“十拿九稳”时,它的自信是发自内心,还是纯属表演。更会揭示,为何总分稳定的模型,答案却可能因一句无关的“废话”就悄悄“叛变”。最后,从给地球装上“大脑”,到揭开我们“猜懂”外语的秘密,这些最新论文将刷新我们对智能的认知。
00:00:33 怎么知道AI“背课文”,而不是“会造句”?
00:06:04 AI的“心里有底”,到底是怎么回事?
00:14:01 AI大模型,总分没变,答案却悄悄“叛变”了
00:18:40 你的地球专属“大脑”,是怎么被训练出来的?
00:25:22 我们其实都是半个翻译家
本期介绍的几篇论文:
[LG] Extractable Memorization From First Principles
[Stanford & Google Research & Google DeepMind]
https://arxiv.org/abs/2607.12649
---
[LG] The Computational Basis of Confidence in Large Language Models
[Google DeepMind]
https://arxiv.org/abs/2607.12447
---
[CL] The Illusion of Robustness: Aggregate Accuracy Hides Prediction Flips under Task-Irrelevant Context
[Georgia Tech & Stanford University]
https://arxiv.org/abs/2607.12963
---
[AI] The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning
[Google Public Sector]
https://arxiv.org/abs/2607.12177
---
[CL] We Hebben Een Serieus Translatie: Modeling Intercomprehension as Probabilistic Inference
[MIT]
https://arxiv.org/abs/2607.12169