推荐 ·

【翻译】How to Be Good at Research:如何擅长做研究

一篇关于研究能力的中英对照翻译:如何选择问题、训练研究品味、升级信息输入、写作思考、缩短反馈回路,并在长期主义中积累优势。

原文:https://x.com/itsreallyvivek/status/2064686372737454155

How to be good at research

Nobody really teaches you research. You get a desk, a problem someone else picked, and a vague instruction to produce something novel. So most people reverse-engineer the job from what they can see, which is papers, threads, and announcements, and what they end up learning is how to look like a researcher rather than how to be one. The actual skill is a stack of smaller skills, and almost every one of them can be deliberately trained.

其实没有人真正教你怎么做研究。你得到一张桌子、一个别人替你选好的问题,以及一句模糊的指令:做出一点新东西。于是大多数人只能从自己看得见的东西里反推这份工作,比如论文、帖子串和项目公告。最后他们学会的,往往是如何看起来像一个研究者,而不是如何真正成为一个研究者。真正的研究能力,是由一组更小的能力堆叠起来的,而其中几乎每一项都可以被刻意训练。

Pick your own problems

Richard Hamming had a habit at Bell Labs that made him unpopular at lunch. He’d ask whoever sat near him what the important problems in their field were, then ask why they weren’t working on them. People changed tables. The question stings because most of us have no good answer. We don’t choose problems, we absorb them, from an advisor, from whatever a big lab announced last quarter, from the paper everyone is quote-tweeting this week.

Richard Hamming 在贝尔实验室有个习惯,让他在午餐时间不太受欢迎。他会问坐在旁边的人:你所在领域最重要的问题是什么?然后接着问:那你为什么不在做这些问题?后来大家会换桌子坐。这个问题刺痛人,是因为我们大多数人都没有一个好答案。我们并不真正选择问题,而是在吸收问题:从导师那里,从某个大实验室上季度发布的方向那里,从这周所有人都在转发评论的论文那里。

The trouble with an absorbed problem is that you hold the conclusion without the reasoning. You know some famous lab cares about a direction, but not why, not what they expect to find, not what would make them drop it. When they pivot, you find out a year later. And on a problem that’s already fashionable, you’re racing a thousand people who started earlier and have more compute than you.

被吸收来的问题有个麻烦:你拿到了结论,却没有拿到推理过程。你知道某个著名实验室在意某个方向,但你不知道为什么,不知道他们期待发现什么,也不知道什么证据会让他们放弃这个方向。等他们转向时,你可能一年后才知道。而在一个已经流行起来的问题上,你是在和一千个更早出发、算力也比你更多的人赛跑。

John Schulman’s guide to ML research splits the work into two modes. In one, you read the literature and hunt for things to improve. In the other, you choose an outcome you genuinely want to exist and reason backwards to the experiments. He argues for the second, and the quiet reason is that it manufactures originality. A goal you actually care about will drag you into territory no survey paper covers.

John Schulman 的机器学习研究指南把研究分成两种模式。一种是阅读文献,然后寻找可以改进的地方。另一种是先选择一个你真心希望它存在的结果,再从这个结果反推需要做哪些实验。他更赞成第二种。背后那个不那么张扬的理由是:这种方式会制造原创性。一个你真正关心的目标,会把你拖进没有任何综述论文覆盖过的地带。

Taste, meanwhile, gets discussed like a gift. It behaves more like a muscle. Predict the result of every experiment before you run it. Cover a paper’s results section and guess the numbers from the method alone. Mark down which of this month’s releases will matter in two years and check your hit rate later. A forecast plus a correction, repeated a few hundred times, is how every good model gets trained, including the one in your head.

与此同时,研究品味常常被讨论得像是一种天赋。但它更像肌肉。每次实验开始前,先预测结果。读论文时遮住结果部分,只根据方法猜测数字。记下这个月发布的成果里,哪些两年后仍然重要,然后以后回来检查自己的命中率。一次预测加一次修正,重复几百次,所有好模型都是这样训练出来的,包括你脑子里的那个模型。

Upgrade your inputs

Shared reading lists produce shared ideas. If your information diet is the trending page of arXiv plus whatever survives the group chat filter, you will reliably reach the same conclusions as everyone else, at the same time, which makes those conclusions worth approximately nothing.

共享的阅读列表会生产共享的想法。如果你的信息饮食只是 arXiv 热榜,再加上群聊过滤后剩下的那点东西,那你会非常稳定地和所有人在同一时间得出同样的结论,而这些结论的价值大约等于零。

Old material is criminally underpriced. This field reruns its own past on a delay: mixture of experts dates to 1991, LSTMs to 1997, backprop went mainstream in 1986. Rich Sutton needed about a thousand words in 2019 to write The Bitter Lesson, and it predicts the shape of the field better than surveys ten times its length. Claude Shannon gave a talk on creative thinking in 1952 where his opening move was to shrink a problem until it’s nearly trivial, crack the small version, then reintroduce the difficulty one piece at a time. That single trick will carry you through more walls than any modern productivity advice.

旧材料被严重低估了。这个领域总是在延迟重演自己的过去:混合专家模型可以追溯到 1991 年,LSTM 是 1997 年,反向传播在 1986 年走向主流。Rich Sutton 在 2019 年只用大约一千词写下《苦涩的教训》,却比长度十倍于它的综述更能预测这个领域的形状。Claude Shannon 在 1952 年做过一次关于创造性思维的演讲,他开场的方法是:把一个问题缩小到几乎微不足道,先破解这个小版本,然后再一块一块地把难度加回来。这个单一技巧,会比任何现代生产力建议都更能帮你穿过障碍。

Range matters as much as depth. Interpretability borrows shamelessly from neuroscience. Eval design is mechanism design wearing a lab coat. A working sense of how GPUs actually move memory tells you which architecture papers are doomed before the benchmarks do. And honest statistics might be the rarest skill in ML, where a lot of published rigor is vibes with error bars.

广度和深度一样重要。可解释性研究毫不客气地从神经科学借东西。评测设计就是穿着白大褂的机制设计。只要你对 GPU 实际如何搬运内存有工作层面的理解,就能在 benchmark 之前看出哪些架构论文注定行不通。而诚实的统计学也许是机器学习里最稀缺的技能,因为很多已发表的严谨性,不过是带着误差棒的感觉。

One more thing. Read the paper itself, not the thread summarizing it. The appendix is where the bodies are buried, and the limitations section is usually the most honest paragraph in the document.

还有一件事。读论文原文,不要只读总结它的帖子串。附录通常是埋尸体的地方,而局限性部分往往是整篇文档里最诚实的一段。

Write everything down

Paul Graham points out that an idea can feel fully formed right up until you try to put it into words. The page finds gaps your head papers over: the assumption you never tested, the step that doesn’t actually follow, the two claims that quietly contradict each other.

Paul Graham 指出,一个想法在你试图把它写成文字之前,可能一直感觉已经完全成形。页面会发现你脑子自动糊过去的裂缝:那个你从未测试过的假设,那个其实并不成立的推导步骤,那两个悄悄互相矛盾的主张。

Feynman’s rule was that the first person you must avoid fooling is yourself, because you’re the easiest target. Writing is the cheapest defense ever invented. Darwin went further and made it procedural. Any fact that cut against his theory got written down on the spot, because he’d caught his own memory deleting inconvenient evidence faster than the convenient kind. Your memory does the same thing to your failed runs. Keep a log: hypothesis, setup, expectation, result, updated belief. Rereading last month’s entries is humbling in a way no reviewer can match.

Feynman 的规则是:你首先必须避免欺骗的人就是你自己,因为你是最容易被骗的目标。写作是人类发明过的最便宜的防御方式。Darwin 走得更远,他把这件事程序化了。任何和他的理论相冲突的事实,他都会当场写下来,因为他发现自己的记忆删除不方便证据的速度,比删除方便证据更快。你的记忆对失败实验也会做同样的事。保持日志:假设、设置、预期、结果、更新后的信念。重读上个月的记录,会带来一种任何审稿人都无法匹敌的谦卑感。

Then put some of it in public. Olah and Carter’s research debt essay makes the case that fields choke on undigested ideas, and that a clear explanation is a genuine contribution rather than a service job. A lot of people working in interpretability today found the field through readable posts, not conference papers. A body of public writing also doubles as the strongest credential you can hold, because it’s an unfakeable sample of how you think.

然后,把其中一部分公开出来。Olah 和 Carter 关于研究债务的文章指出,领域会被未被消化的想法堵住,而清晰解释本身是真正的贡献,不是服务性工作。今天很多做可解释性研究的人,最初是通过易读的文章进入这个领域的,而不是通过会议论文。一批公开写作也会成为你能持有的最强凭证,因为它是你思考方式无法伪造的样本。

Tighten the loop

The stories about Alec Radford rarely involve a single stroke of genius. They involve volume. More runs per day, more wrong ideas discarded per week, a model of reality that updated faster than anyone else’s. That’s the actual game. Research speed is mostly the speed at which you discover you’re wrong.

关于 Alec Radford 的故事,很少是某一次天才般的灵光乍现。它们更多关乎数量。每天更多次实验,每周丢掉更多错误想法,一个比别人更新得更快的现实模型。这才是真正的游戏。研究速度,本质上大多是你发现自己错了的速度。

Which makes tooling a first-class research activity. Launching a run should be one command. Plotting it should be one more. Every experiment should be reproducible from its config, and comparing two runs should take seconds, not an afternoon of archaeology. Karpathy’s recipe for training neural networks has a step that pays for itself a hundred times over: overfit a single batch before training at scale. Thirty seconds, half your bugs, gone. Shrink everything until it’s cheap, get it right, then spend the compute.

这使得工具建设成为一等研究活动。启动一次运行应该只需要一条命令。画图也应该只需要再一条命令。每个实验都应该能从配置文件复现,比较两次运行应该只花几秒钟,而不是一个下午的考古。Karpathy 的神经网络训练配方里有一步,回报率高得惊人:在大规模训练前,先过拟合一个小 batch。三十秒,半数 bug 消失。把一切缩小到便宜为止,先把它做对,然后再花算力。

And retire the idea that engineering is the junior partner here. At the frontier the two jobs have fused. The researcher who can build the harness, the eval, and the data pipeline is the one whose hypotheses actually get tested. Everyone else is waiting in a queue.

同时,也该放弃“工程只是研究的低阶伙伴”这种想法。在前沿位置,这两份工作已经融合了。能搭建实验框架、评测和数据管线的研究者,才是那个假设真正能被测试的人。其他人都在队列里等待。

Stare at the outputs

A descending loss curve is not analysis, it’s reassurance. Your experiments throw off far more information than you consume: transcripts, failure cases, the strange tail of the distribution. Most of it dies unread in a logs folder.

下降的 loss 曲线不是分析,它只是安慰。你的实验产生的信息远远多于你实际消费的信息:对话记录、失败案例、分布里奇怪的长尾。它们大多数都死在无人阅读的日志文件夹里。

Karpathy’s recipe starts before any training code gets written, with hours spent on the raw data by hand. Most ML bugs live in the data, and they fail silently. Nothing crashes. You simply get a mediocre model and a wrong theory about why.

Karpathy 的配方在任何训练代码写出来之前就开始了:先花几个小时手动看原始数据。大多数机器学习 bug 都藏在数据里,而且它们会静默失败。没有东西崩溃。你只是得到一个平庸的模型,然后对为什么平庸形成一个错误理论。

Andrew Ng has taught the same unglamorous move for over a decade because nothing beats it. Pull a hundred failures, read all of them, sort them into piles, attack the biggest pile. It works on models and it works on evals, where a benchmark you’ve never read transcripts from is a benchmark you don’t actually understand. One transcript of genuinely strange behavior will teach you more than the next decimal of accuracy ever will.

Andrew Ng 十多年来一直在教同一个不光鲜的方法,因为没有什么比它更有效。抽出一百个失败案例,全部读完,把它们分堆,然后攻击最大的那一堆。这对模型有效,对评测也有效。一个你从未读过转录内容的 benchmark,其实就是一个你并不真正理解的 benchmark。一条真正奇怪行为的记录,教给你的东西会比准确率小数点后再多一位更多。

Wander on purpose

Your first subfield is an accident of timing, so treat it like one. Spend real time in interpretability, in evals, in RL, in systems, before deciding where you live. Somewhere in this field is a corner where your specific weirdness is an unfair advantage, and the only way to locate it is to pay tuition in several places. Nobody waives the tuition.

你的第一个子领域只是时间偶然性的产物,所以也应该这样对待它。在决定自己长期待在哪里之前,真正花时间去可解释性、评测、强化学习、系统等方向里走一走。这个领域的某个角落,一定存在一个地方,你身上某种具体的怪异会变成不公平优势。而找到它的唯一办法,是在好几个地方都交一点学费。没人会替你免掉这笔学费。

Run the disposable version of every idea first and let most of them die young. Tune your baselines until it hurts, because the graveyard of ML is full of gains that evaporated against a properly tuned baseline, and a reviewer is the worst possible person to learn that from. Ablate until you know which component carries the result. It’s usually one, and it’s usually not the one in the title.

每个想法都先跑一个一次性版本,并允许它们大多数早死。把 baseline 调到让你痛苦为止,因为机器学习的坟场里堆满了那些在合适调参的 baseline 面前蒸发掉的增益,而审稿人是最糟糕的提醒你这件事的人。一直做消融,直到你知道到底是哪一个组件支撑了结果。通常只有一个,而且通常不是标题里写的那个。

Breadth is also insurance. Subfields saturate, all of them, usually right after they peak on Twitter. The people who keep producing through those transitions are the ones who already know their way around the neighboring territory.

广度也是一种保险。所有子领域都会饱和,而且通常就在它们在 Twitter 上达到峰值之后。那些能在转折期继续产出的人,是早就熟悉邻近地带的人。

Find your people

Hamming noticed a pattern in who ended up doing important work. Colleagues with closed office doors got more done in any given year, and colleagues with open doors did the work that mattered, because the interruptions carried information about what the world actually needed. Your open door is probably an inbox. Keep it that way.

Hamming 注意到,最终做出重要工作的人有一种模式。办公室门关着的同事,在任何一年里都会完成更多事情;办公室门开着的同事,则会做出更重要的工作,因为那些打扰携带了关于世界真正需要什么的信息。你的“开着的门”很可能就是收件箱。让它继续开着。

Generosity compounds in research like nothing else. Replicate a result and publish what you find. Release the tool you built for yourself. Explain something hard in plain language. The returns arrive sideways, months later, as the collaboration or the reference or the role you couldn’t have applied for. Float your half-formed ideas in public too, because being wrong on the timeline is far cheaper than being wrong in print. And the collaborator who tells you an idea is bad before you sink three months into it is worth more than compute. That relationship can’t be bought, only earned.

慷慨在研究中会以其他东西难以相比的方式复利。复现一个结果,然后发布你的发现。公开你为自己做的工具。用清楚的语言解释困难的东西。回报会在几个月后从侧面到来,可能是一段合作、一次引用,或者一个你原本无法申请到的职位。也把你那些半成形的想法放到公共空间里,因为在时间线上犯错,远比在正式出版物里犯错便宜。那个在你投入三个月之前就告诉你某个想法很糟糕的合作者,比算力更有价值。这样的关系买不到,只能赢得。

The long game

Pasteur said luck favors the prepared mind, and Hamming built a whole career philosophy on top of it: knowledge and productivity compound like interest. The daily edges look trivial in isolation. What you read, what you record, how fast your loop runs, who you argue with. Give them a few years and they produce careers that look like luck from the outside. Start compounding earlier than feels necessary. Future you already knows this was the cheap part.

Pasteur 说,幸运眷顾有准备的头脑。Hamming 在这句话之上建立了一整套职业哲学:知识和生产力会像利息一样复利。日常里的微小优势单独看都很不起眼。你读什么、记录什么、反馈回路跑得多快、和谁争论。给它们几年时间,它们就会制造出从外面看起来像运气一样的职业轨迹。比你觉得必要的时候更早开始复利。未来的你已经知道,这一段才是最便宜的部分。