My PhD Grind (1): All the Friends Who Worked on Theory Have Switched to LLMs / 我的PhD Grind(1):23年申phd时认识的那群朋友至此都转了LLM
(The original version by the author is in Chinese without using any LLM, but for English version, the author use Claude to do the translation and then manually edited and rewrite some of them that she thought inaccurate.)
(Note: In order to preserve the privacy of the people that I mentioned, I didn’t use their real name, but instead, using the name “Jack”, “Frank”, “Leo” instead.)
I had originally wanted to give this the structure of a perfect book: a prelude, three distinct chapters (not for the sake of literary completeness, but because my PhD really did come in stages, almost neatly divisible by year), and an epilogue. So after jotting down some loose notes on the flight back to China, I spent a long time thinking about how to structure what I’d written, how to shape it into the perfect, complete form I had in mind before presenting it to readers. The result: three weeks went by, and the draft I wrote on the plane was never finished.
One day in September, while chatting with a friend, I was surprised to learn that he had switched from theory to LLMs. A rush of feelings came over me and I wanted to write something. What I wanted to write belonged, properly speaking, in my PhD memoir—but the memoir had never gotten off the ground, precisely because of my pursuit of perfection. I don’t know how I picked up this habit, but whenever I present something imperfect to readers, I feel I have to apologize first and honestly list everything I think is wrong with it. (Seen this way, I’m really not cut out for academia or startups. Though the people and events of my PhD have also taught me to be a little less harsh on myself, and a little more forgiving.) So I’ll just write freely, wherever my thoughts happen to go. Reality is, after all, imperfect and incomplete; if I tried to write this like a work of literature, with the reader’s experience in mind, I would be sacrificing truthfulness and spontaneity—the feelings I have while writing, and the flow of my state of mind. Besides, works of literature have storylines, and reality doesn’t. I’ve become more and more able to accept the chaos of the world, and my own chaos within it (the flattering term for which is “free-spirited”).
That day, I happened to open an article on Zhihu and noticed that the account of a friend, Leo, had liked it. So I messaged Leo on WeChat—we hadn’t seen each other in a long time, and I wanted to catch up and ask how he’d been. He told me he was planning to switch to working on LLMs, and that over the past year he had moved to a new place and was talking to people even less than before. Even though we hadn’t been in touch for over a year, chatting with him still felt familiar. Back in 2023, when we were applying for PhDs, we were both working on theory. But Leo socialized very little and rarely reached out to anyone to hang out (he spent most of his time doing research and reading papers), though when we ran into each other now and then I’d chat with him a bit. So I wouldn’t say I knew him well enough to call us “close.” And yet he came to my PhD A exam last year. (Not many people came since I hadn’t made any special announcement—it was only visible in a mass email that the grad school requires us to send to all the PhD student in a certain format. He’s also not the kind of person who likes going to presentations, especially when the research I was presenting had nothing whatsoever to do with what he was interested in. So the fact that he noticed, and made time to show his support, surprised me.) Leo’s PhD path has been extremely rocky. Given his hermit-like style, he probably wouldn’t want me telling other people too much about it, so I’ll leave that part out. But even though his PhD hasn’t gone smoothly, he has always been so calm, so unruffled.
Every time I talk with people around me (especially the guys) and learn what they’ve been through, I think: “Why am I so bad at enduring hardship? The smallest thing happens to me (compared to what they have been through) and I start kicking up a fuss, as if I want the whole world to know how much pain I’m in and how hard things are—while they, facing far more and far harder things, keep a straight face, placid as still water.” Of course, this kind of self-gaslighting usually passes in an instant. I tend to look at things from an all-around, complete perspective, so everything gets considered from multiple sides and angles, and I freely adjust the weight I give to each. The upside is that I never become too stubborn or opinionated; the downside is that sometimes people feel I have no personality of my own, that I just agree with everything. (I’ve just wandered off topic again…)
In fall 2024, I was considering quitting, and Leo was in the middle of switching and finding new advisors. I half-jokingly said, “Why don’t you quit with me? You wouldn’t even have to worry, like I do, about how quitting would negatively affect your advisor’s reputation. Or, stop doing theory and switch to an advisor who does applied work?” (In a CS department, if Leo didn’t care about his research direction, finding an advisor would be much easier.) But he remained very firm, and firmly described to me the career path he had in mind. He framed his choice to do a PhD as “so that I can go into quant after graduation and get a very high salary.” I didn’t quite follow his logic, because from my perspective, he could go into quant companies even without a PhD, or without doing theory. But Leo described it as though it were the only path. I just smiled, because from what I knew of him, he was the unpretentious kind of scholar, and the “high salary” of quant would hardly be enough to make him choose, so resolutely, a path that looked so much harder. (Note that I added the qualifier “unpretentious” here, because not every scholar can be described that way—and this was precisely the mistake I made when I first entered academia. But nothing is absolute; back then I could only see one side of people and things, and I blindly passed judgment on them based on what I saw. That was my own problem.)
Perhaps it’s Leo’s “hermit” style that lets him hear less of the world’s noise, and steadfastly do what he likes, what makes him happy (doing theory, writing formulas, and reading proofs really are a joy for some people!). In my first year of the PhD, I too agonized for a long time over whether to give up theory. Even back when I started doing theory as an undergrad, my undergrad advisor kept telling me that doing theoretical research has a lot of downsides, such as it is hard to recruit students, hard to get funding, and students are hard to find a job… But at that time, working are theory made me feel calm and happy, so none of those practical considerations mattered. I had hoped I could do both theory and application (so in my first year I was in fact doing both at once). On the one hand, I couldn’t give up theory, because it made me feel at peace; on the other hand, I couldn’t give up NLP, because coding had been my weak spot for the past several years, and if I abandoned applied work entirely, it would become my weak spot once again. (And with that would come a lack of confidence, the kind that keeps me from seeing many people and things from an objective, equal footing, but admire people who can code.) I struggled for an entire year since I don’t want to give up any of these two, and in the end, because of a project I didn’t much like, I finally decided to stop working on theory. Looking back now, I ought to thank this twist of fate, because I was no longer distracted the way I used to be. Before, even when doing theory and applied projects side by side, I would always prioritize theory, because it made me happy. I didn’t care whether I was doing it well enough, didn’t care that a simple proof took me ages to understand, didn’t care whether my brain or my level of effort were suited to it. I only cared about the happiness of the moment, and about that slow, unhurried life of “one pen, one book, one proof to stare at for days.” (Even though some of the peers I know, could probably get through it in 30 minutes, lol.)
When my second year began and I decided to give up theory completely and focus on LLMs and NLP, it was painful at the time, because I didn’t like living at such a fast pace, and I didn’t like writing code. I wanted my tools to be pen and paper; I wanted the computer to be something I used for reading proofs, not for writing code and running experiments. I still tried hard to steer clear of the research directions that looked very popular, but I really could no longer get the joy of writing proofs, nor the calm I’d had before. Now, two years later, I’ve grown used to my current identity and come to accept it—I can’t even understand anymore why my past self was so afraid of giving up theory, so afraid of working on a popular direction. Back then I simply just being too young too naive; every little thing seemed like a big deal, and I didn’t dare try anything outside of what I had mentally prepared myself for. Now I’ve become more nihilistic: once there are more things you don’t care about, you’re less afraid of losing them. And once you’ve seen enough data, you know that people can always adapt, little by little; you don’t need to keep yourself in the most comfortable environment all the time, and you don’t need to confine yourself to a box. (Writing blogs is the new thing that makes me feel calm. Writing my thoughts and experience down is actually quite a lot like writing proofs, except that when I wrote a proof, I felt I was presenting a perfect piece of work, and the recognition of those around me made me sure of it; when I write prose, my thoughts often drift far off, so, objectively speaking, I can’t produce a piece of work that would be called good by any objective standard. But one day, I suddenly realized that the people and the world around me are more accepting than I am to my work—things don’t have to be perfect before you can show them to others. (I tried a few more times after that and received a lot of positive feedback.) I didn’t post these writings on English social media; but I quietly put them on my homepage. (For Chinese versions I occasionally post on Chinese social media, hoping strangers with no connection to my life will read them) One day, I received an email in from a stranger who doesn’t speak Chinese, saying she really liked my blog, and that she can relate to what I write even though we have a quite different life trajectory. This kind of good moments are one of the reasons why I like writing: it is for in this enormous world, to briefly form a super weak connection with a stranger that you would have never had any relation with (i.e., form an edge in two atom in a world full of atoms), and then go back to being strangers. It was also because of this email that I later decided to proofread my English versions carefully in my future writings, instead of simply translating them from Chinese without further proof-reading.)
At ICML 2024, even though I had already decided not to working on theory and spent most of my time talking to peers working on ML applications, three people who works on theory, whom I know when we were all applying PhD in 2023 happened to be at the conference too. One of them, Frank, invited us all for a city walk to catch up with each out, since we have each others contact but have never met in person. Being with them, I felt a bit out of place, because by then I could no longer follow the optimization theory and stochastic algorithms they were so heatedly discussing. I could only listen, envying them for still pursuing what they loved, even with so many LLM chaos all around. (Frank probably invited me along because, if when you grabbed a random person at ICML back then, the random person were most likely working on LLMs or something related. So unlike those who works on LLMs, who should be busy networking and talking with people in their own field—they could only grab people that they already know to catch up with.)
At NeurIPS 2025, two of those three person I mentioned about attended. One had already switched to LLMs and was busy talking with other LLM people at the conference—too busy to hang out with us. Thus, Frank and me decided to catch up ourselves. He said that he and the third guy who l used to work on theory (the person who didn’t attend NeurIPS) were both preparing to switch to LLMs (Frank and I each conveyed to him with our most sincere and heartfelt longing through messages), and we also discussed things like switching research directions and finding internships. Because NeurIPS was in San Diego, a huge number of people came, and I also ran into a younger PhD student Jack. He too had switched from theory to LLMs, and was also very anxious. We discussed the state of the job market. By then there were already many PhDs graduating in only three or four years (I was job hunting at the time), and I was anxious in thinking that: if all those PhD students, who should have entered the market one or two years later suddenly flood in all at once, doesn’t it get even harder for not-so-outstanding PhDs like me, or for people who’ve only just switched their research directions, to find jobs? (On top of that, there were also many extremely talented first or second year PhDs who leave PhD program to join industry.) Jack, anxious but forcing himself to look on the bright side (he had only just started his second year), said, “Or maybe I’ll just stay in school through this turbulent period, and by the time I graduate the market will have stabilized again.”
I also chatted with a peer I had once looked up to—a rising star in my eyes, the one who “could read in 30 minutes a proof that took me days to understand.” He said he had switched to LLMs in the past few months, but since others had several more years of experience than he did, he was slowly feeling his way forward, trying to catch up. Once again I didn’t know what to say. I thought: even those who once shone so brightly in their own field can feel lost and adrift when faced with something new. If he had stayed in his field, what a dazzling star he would have been. But LLMs have forced many people to abandon the directions they love and are good at, to work on something new, while others who got into it a little earlier already have a lot of experience. Sometimes I wonder: is it possible to catch up? Judging by most people’s examples, it seems that as long as you’re willing to work hard, willing to put in the time, and quick to grasp things, catching up is possible—and I’ve since seen many examples of successful transitions. Though I’m not one of the “willing to work hard, willing to put in the time, quick to grasp things” crowd (by now I’ve made peace with my own nature, so I just take things as they come and have no plans to catch up with anyone). But the those people I mentioned above—those who used to work on theory, they’ll surely catch up in the end, through their hard work and their brilliance.
By 2026, almost every person met while applying for PhDs who used to work on theory had switched to LLMs. The trigger for writing this piece was that the friend Leo I mentioned earlier, the one who had wanted to go into quant, told me today that he just switched to LLMs this semester—he who had once been so resolute even in his most difficult time when he can’t find an advisor. He said he had only just changed directions, was still in the adjustment period, and was very unhappy due to the transitioning phase.
As I write this, my mind is flooded with thoughts. Our PhDs—a mere three years, and yet it feels like such a long, long time has passed. So long that, unless I make an effort to remember, I’ll forget that many of those people and ephemeral moments ever existed. So long that when I sit down to write, I don’t know where to begin, or how to end. But these people and moments are still worth remembering. One of the reasons I originally wanted to work on ML theory was that I felt everyone in this fields was so nice (after all, everyone had plenty of free time, did research at an unhurried pace, and was very easygoing)—whether Leo, who didn’t usually hang out with me but showed up at my A exam to show his support, or Frank, whom I only knew each other online due to PhD applications, but were thoughtful enough to find a reason to hang out everytime when he know that we are attending the same conference. Our friendships feel so light and understated—so light that I often forget them constantly, and they might also forget me. But when something reminds us of each other, we can still catch up like old times. Even though so much has changed—we’ve all switched directions, and life and work have made us less idealistic than we were—there’s always something familiar that hasn’t changed. When we grow older, we stopped chatting with people without a specific purpose. Everyone has gotten into the habit of talking with a purpose (“with a purpose” isn’t meant as a criticism here; it’s more that a conversation needs a reason, and most of the time my reason is work or academic related). So I don’t know if we are friends or coworkers, since everying conversation feels like “meeting” or business chat.
Those friendships with those old friends (who used to work on theory), though seem so light, yet don’t feel the least bit unfamiliar when picked up again a year later, are pure enough to leave me nostalgic and savoring.