Checkpoint One

Jul 3, 2026

I'm one year into the PhD. I've been watching people I went to school with scatter in directions I know so little about: friends playing professional soccer in Europe, teaching elementary school, grinding away at AI startups (whom I check on occasionally to confirm they sleep), collecting clinical hours while working two jobs, traveling the world and WWOOFing on farms, settling in with partners and starting families, and some (well, me) spending most days talking to chatbots.

Every once in a while people ask me if I'm glad I'm in a PhD program. Whether I made the “right” decision. Recently I came across a line from Professor Ellen Langer: don't make the right decision; make the decision right.

So the useful question isn't whether the PhD was the right choice. It's how to make it a valuable, challenging, growth-inducing, dare I say magical, time.*

Hence this audit of my first year. Being so deep in Claude lately, I also wanted a creative exercise, one that forced me to sit with a healthy amount of writer's block without reaching for a chat window. Bear with the over-intellectualization: what follows is a self-reflection that casually plays with the parallels to an agent's framework.

Others have done the human-as-algorithm metaphor properly. Brian Christian's Algorithms to Live By takes the decision-making side, when to explore versus exploit; Rohan Kshirsagar's Why viewing yourself as a black box is one of the best things you can do takes the self-modeling side. I absorbed the broadest version, life as an optimization problem, least deliberately, through Berkeley's ML courses. When I was weighing post-grad plans, a professor (now one of my advisors) calmly delivered: “Just do gradient descent.” A year in, consider this checkpoint one.

* I have on occasion been accused of toxic positivity.

On Environment

I got deployed out of distribution this year. Undergrad was more or less a stationary environment: the semesters had a shape, and I had a long list of productive tasks to complete. Research doesn't hold still. The field moves every few weeks, the norms differ from lab to lab, and the terrain you mapped in the fall is partially fiction by spring.

Meanwhile, academia seems to be sprinting to build the “next generation” of evaluation environments: more prompts to test, more red-teaming, and more domains in which we must try agents. Lately there's a surge of non-stationary environments (sound familiar?) that shift underfoot and demand continual learning. The field isn't obsessed with evaluations because it ran out of other ideas. Environments simply control agent behavior more than we may initially realize.

People are no exception. My parents always told my siblings and me to be deliberate about the humans we let into our lives, because their values will wear off on you. Float a slightly risky opinion at dinner and get a beat of silence and three blank stares. That's feedback, and you can feel your next sentence already softening. Make an offhand joke, the table actually laughs, and some part of you files it away to run again. I come home to my roommate eating a mango and coincidentally buy one later that week. None of this rose to the level of conscious choice.

That sort of thing unsettles me about real human environments: the updates happen whether or not we're aware of them. If I can't opt out of the training, I can only try to control what's around me. And while I don't get to pick all of my input data, I get more say than I thought.

So a lot of my first year was, in retrospect, an attempt to work that margin. I found advisors whose taste in problems I'd be glad to sample. I learned that I think differently on the sixth-floor Soda balcony than I do on the couch with a cup of tea (and that's a fact about how I work rather than a preference). I put real effort into who I surround myself with, choosing whose blank stares and laughs get to shape me. At times, this effort felt like unproductive, frustrating logistics that just needed to be done. But sitting with it longer, that's a core assignment of grad school: attempting to build an environment that serves me well has demanded a level of intention I'd never exercised before.

On Architecture

Aside from where an agent is deployed, the next big design decision is which model to choose. Within language models, the choice is between open-weight and black-box: one lets you adapt the weights directly (post-training, RL, etc.); the other you can't touch from the inside. You can only steer it by changing the prompt and the harness.

Suppose I'm an open-weight model. The point of open weights was never being able to interpret them; it's being able to change them, a little per update. That's how learning to ride a bike went: every attempt, my feet found the pedals a bit faster, my knuckles less white, and nothing I could name was different. I talked myself through it at first (“look ahead, not down”), but that self-talk was scaffolding; it dropped away and the skill stayed, settled into weights I've never consciously touched. This is the belief behind deliberate practice: run the reps, trust the update. Under this theory, if I just read enough textbooks and papers closely, I must be learning something in my weights, even if I can't say exactly what it is or how it connects to my current research project.

Suppose instead I'm black-box. Then I've given up on reaching the weights directly, and I work the inputs instead: the framing, the harness, the prompt I feed myself. And a lot of what I feed myself is other people, what they believe, how they frame a problem, which parts they think matter. Much of black-box learning, it turns out, is consolidating those outside views into a mental framework of my own. It's messier, but it forces me to make my hypotheses about my own learning explicit. This might be a real advantage. (This is more or less Kshirsagar's argument, and it's the version I trust most on the days when willing myself to change has obviously stopped working: when no amount of “just try harder” moves anything, but changing the framing does.) Here, when I read the textbook, I have to pause and write out the key takeaways and connections to my research in a format I can reference back later.

Is one mental model of learning necessarily better than the other? I don't think so. They're two different tools, and the skill is figuring out which one fits the moment: grinding away at deliberate retraining on something that will only shift if you reframe the problem, or waiting around for context to fix a thing you could just practice. When I'm stuck, I'm still learning to make a diagnosis: am I just under-trained or do I need a better framework?

On Training Dynamics

Push the learning rate too hard for too long, and you get dead neurons, units that have simply stopped firing. Maybe metaphorically, burnout is just a bunch of dead ReLUs.

Cortisol is something like a learning rate scheduler. Useful in short bursts (real threat, real deadline), but if it never comes back down, you don't learn faster. You learn worse. Attention narrows, memory consolidation suffers, and you start reacting to threat instead of updating on new information. The scheduler fails in the other direction too: in stretches with no deadlines at all, nothing really moves, and somehow no urgency means no updates either. I don't do my best thinking in weeks when everything feels urgent, but I don't do it in weeks when nothing is either. I do my best work when something has had time to sit, with just enough pressure that it can't sit forever.

Which gets to the other half of training dynamics I keep circling: not just whether the system has crashed, but how much evidence should actually move it? Not every off day means my approach is wrong, and not every good day means I've found the answer. But I have a bias toward updating too fast on emotionally loud data points (one unproductive meeting, one hard day) and too slow on quiet, accumulating ones. A single strong gradient shouldn't outweigh a hundred small ones, but it's the one that's the least noisy and easiest to notice.

Part of this is also being willing to hold multiple things that look like contradictions without immediately resolving them, a willingness the last year has tested more than once. Someone complicates what you think, and if your reflex is to defend, you've frozen the weights on purpose. No gradient gets through. Most of us can catch ourselves, on occasion, listening to respond rather than to update our beliefs. Good updating often means sitting with “this is also true” for a while before deciding what to do with it. Academia prides itself on criticism and rigor, so maybe disagreement is an opportunity to get more data; I'm in the process of learning to receive it that way.

On Objectives

This is supposed to be the section where I unpack self-discipline as a designed reward signal. Perhaps I say something wise about habits and finishing that 15-mile run.

The annoying thing about research objectives is that the reward is sparse and delayed almost by design. Nothing pings back when you read a paper closely, or when an experiment fails quietly instead of loudly. A real reward, a result, a paper, is years out, if it comes at all. So you're always running on some invented proxy instead, because the true signal doesn't arrive on a timescale a nervous system can use.

Optimize the proxy instead of the thing it stood in for, and you're reward hacking. I can measure my productivity by “hours in front of Claude” or “papers read,” and neither is the actual objective, just the thing I could measure that week.

My old soccer coach had the low-tech version of this speech: cut corners in training, and you're only cheating yourself. But with AI it's more convoluted than just willing myself to “not cheat,” and it's become a source of anxiety. I'm scared to let AI run experiments for me, and scared not to, because I can't always tell which choice protects the “real” objective. Sometimes I'll have Claude digest a paper and quiz me until I am convinced I understand, and “convincing myself” is exactly the problem: real understanding or a new proxy wearing the old one's clothes?

Sometimes I need to design more proxies into my research process. One advisor suggested a timeline with milestones; another, a research log. Different functions, same advice: stay organized enough to make progress and make that progress visible to myself and others. Sometimes proxies are less dignified. I didn't sit down to write this in order to change my research, or because any sentence along the way provided reward. At some point the proxy was “I said I'd finish this,” which is so far removed from any real research objective, it sounds fake written down. But the fake proxy kept me in the chair, and somewhere in the sitting I received the real reward: my creative brain online and engaged.

In Conclusion

I don't think research, or life for that matter, is just convex optimization. No framing of the problem can perfectly capture everything that matters, and no learning algorithm can promise an optimal solution. The consolation: in high enough dimensions (and the world has plenty), many local minima turn out to be good ones. Which is, I think, what Langer meant: you don't find the right minimum, you make the one you're in right.

So, to the important people in my life scattered across soccer pitches, startups, med schools, farms, and new families (rolling their eyes at me for these silly analogies): it's dizzying and deeply awe-inspiring to watch you all veer in directions I know so little about. There's a specific magic to this age: everyone in motion. I love seeing each of you architect different agents, in different environments, chasing different rewards.