<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://dharakyu.com/feed.xml" rel="self" type="application/atom+xml" /><link href="https://dharakyu.com/" rel="alternate" type="text/html" /><updated>2026-07-28T04:46:52+00:00</updated><id>https://dharakyu.com/feed.xml</id><title type="html">Dhara Yu</title><subtitle>Personal website and blog</subtitle><entry><title type="html">Does social media research hold any useful lessons for studying the effects of AI on human behavior?</title><link href="https://dharakyu.com/blog/2026/07/10/social-media-vs-ai/" rel="alternate" type="text/html" title="Does social media research hold any useful lessons for studying the effects of AI on human behavior?" /><published>2026-07-10T00:00:00+00:00</published><updated>2026-07-10T00:00:00+00:00</updated><id>https://dharakyu.com/blog/2026/07/10/social-media-vs-ai</id><content type="html" xml:base="https://dharakyu.com/blog/2026/07/10/social-media-vs-ai/"><![CDATA[<p>I read a recent correspondence in Nature Human Behavior, provocatively titled <a href="https://www.nature.com/articles/s41562-026-02482-9">“Why social media research has failed policy-makers”</a>, by Tobias Dahl et al.</p>

<p>This essay caught my eye, in part because it is relatively uncommon to see researchers publicly criticize their own fields and research programs (although this is a recent <a href="https://smallpotatoes.paulbloom.net/p/a-lot-of-developmental-psychology-f0a">counterexample</a>). But beyond that, its title invited connections to some of the questions I’ve been thinking about, namely how interactions with AI are changing aspects of human cognition and behavior: there are clear metascientific parallels between this emerging line of work and past / ongoing research that examines how social media use affects wellbeing. To be sure, there are many qualitatively new and different phenomena at play with AI that make it far more than “social media 2.0”. But the methodological challenges we face in studying these questions about AI interaction mirror many of the challenges endured by social media researchers - limited access to proprietary data, ill-defined dependent variables, and difficult-to-resolve tradeoffs about the value of highly-controlled but artificial lab studies vs. more naturalistic yet potentially uninterpretable observational analyses. Is there anything we can learn from where social media research has erred?</p>

<p>Before we get there, let’s unpack the diagnosis provided by Dahl et al. First, they argue that most social media wellbeing research actually doesn’t find a negative relationship between time spent on social media and various metrics of wellbeing. But this does not license the conclusion that social media is benign. Rather, this effect could be driven by the cost of missing out (COMO) phenomenon, where adolescents who are restricted from using social media actually fare worse on wellbeing metrics because they cannot partake in a shared social experience. To make causal claims about the effects of social media, there needs to be a proper control condition in which there was an entire group or community of adolescents not on social media but otherwise matched to the “experimental” group of youth on social media - something that is of course very difficult to come by. Without this setup, we are left with a deeply limited and ultimately non-actionable base of evidence.</p>

<p>I’m not fully convinced by the specific COMO explanation offered here - if one is pre-committed to a particular point of view (e.g. social media is or is not harmful), one could concoct any number of confounds that explain why the result apparently goes in the other direction (see <a href="https://www.theatlantic.com/technology/2026/07/phones-haidt-play-gray/687846/">this</a> for more examples of how this is playing out in the social media debate). I’m generally sympathetic to claims that social media is a net negative to wellbeing, but it’s unclear to me if COMO is actually a reasonable explanation as to why the evidence looks mixed. But more broadly, I think it’s worthwhile to highlight the importance (and difficulty) of having a reasonable control / baseline group to draw causal claims - <a href="https://benmtappin.substack.com/p/are-ai-chatbots-harmful-or-beneficial">a recent essay emphasized this precise point</a> in trying to assess the effects of chatbots in domains such as mental health.</p>

<p>Of course, principles of experimental design are not unique to social media research; this is just one manifestation. So despite my initial optimism that we could learn something from our predecessors, I am left with the impression that little from the social media playbook is actually transferable to AI interaction research, beyond serving as a generic cautionary tale about the difficulty of trying to pursue research when the most illuminating data is proprietary.</p>

<p>To highlight a specific methodological area of disanalogy: this essay seems to suggest that within the social media wellbeing research community, it is a common practice to bundle all interactions with social media into a catch-all “use” or “time spent” quantity. I’m not sure this was ever a great methodological choice, but it seems even less useful to do that when studying the effects of AI, because the set of things one can do with AI is much broader, even compared to social media. Imagine 2 study subjects: one who uses AI to speed through school assignments, versus another who uses AI to create practice problems to then solve independently. Under a coarse-grained “use” metric, these two subjects could very well be classified into the same bucket, which feels obviously wrong. There are inertial forces (e.g. the fact that it’s relatively easy to get self-reports of time spent on a specific AI product) that will likely push researchers toward the catch-all “use” value as a predictor variable, but ultimately that practice does not seem like it should be rolled over.</p>

<p>To me, this case study showcases the limits of analogical thinking about how past technologies have shaped human cognition and behavior. Ultimately, I actually found it a more useful exercise to think about the ways in which AI is <strong>not</strong> like social media, than to try to draw explicit parallels. For a more well-formed example of what this contrastive approach can buy you: we can have more secure foundations to reason about <a href="https://www.conspicuouscognition.com/p/how-ai-will-reshape-public-opinion">how LLMs could reshape public opinion</a>, by contrasting it with the effects of social media.</p>

<p>If we hope to arrive at a holistic picture of how AI is changing how we think, there are a lot of hard problems that need to be addressed; social media research didn’t pave the way methodologically, although maybe it offers a useful point of contrast. An additional consolation is that the object of study is so flexible as to be a useful analytical tool to help us answer these questions, more quickly: AI, used judiciously, could potentially help unblock some of these methodological challenges (e.g. coding agents could make it easy to build a more granular tracker of individuals’ interactions with AI). The big question remains, though, of how we will sift through an ever-expanding body of AI-powered research to understand what this technology is doing to our own minds - and if we can draw conclusions at the speed at which AI is infusing itself into our lives.</p>]]></content><author><name></name></author><summary type="html"><![CDATA[I read a recent correspondence in Nature Human Behavior, provocatively titled “Why social media research has failed policy-makers”, by Tobias Dahl et al. This essay caught my eye, in part because it is relatively uncommon to see researchers publicly criticize their own fields and research programs (although this is a recent counterexample). But beyond that, its title invited connections to some of the questions I’ve been thinking about, namely how interactions with AI are changing aspects of human cognition and behavior: there are clear metascientific parallels between this emerging line of work and past / ongoing research that examines how social media use affects wellbeing. To be sure, there are many qualitatively new and different phenomena at play with AI that make it far more than “social media 2.0”. But the methodological challenges we face in studying these questions about AI interaction mirror many of the challenges endured by social media researchers - limited access to proprietary data, ill-defined dependent variables, and difficult-to-resolve tradeoffs about the value of highly-controlled but artificial lab studies vs. more naturalistic yet potentially uninterpretable observational analyses. Is there anything we can learn from where social media research has erred? Before we get there, let’s unpack the diagnosis provided by Dahl et al. First, they argue that most social media wellbeing research actually doesn’t find a negative relationship between time spent on social media and various metrics of wellbeing. But this does not license the conclusion that social media is benign. Rather, this effect could be driven by the cost of missing out (COMO) phenomenon, where adolescents who are restricted from using social media actually fare worse on wellbeing metrics because they cannot partake in a shared social experience. To make causal claims about the effects of social media, there needs to be a proper control condition in which there was an entire group or community of adolescents not on social media but otherwise matched to the “experimental” group of youth on social media - something that is of course very difficult to come by. Without this setup, we are left with a deeply limited and ultimately non-actionable base of evidence. I’m not fully convinced by the specific COMO explanation offered here - if one is pre-committed to a particular point of view (e.g. social media is or is not harmful), one could concoct any number of confounds that explain why the result apparently goes in the other direction (see this for more examples of how this is playing out in the social media debate). I’m generally sympathetic to claims that social media is a net negative to wellbeing, but it’s unclear to me if COMO is actually a reasonable explanation as to why the evidence looks mixed. But more broadly, I think it’s worthwhile to highlight the importance (and difficulty) of having a reasonable control / baseline group to draw causal claims - a recent essay emphasized this precise point in trying to assess the effects of chatbots in domains such as mental health. Of course, principles of experimental design are not unique to social media research; this is just one manifestation. So despite my initial optimism that we could learn something from our predecessors, I am left with the impression that little from the social media playbook is actually transferable to AI interaction research, beyond serving as a generic cautionary tale about the difficulty of trying to pursue research when the most illuminating data is proprietary. To highlight a specific methodological area of disanalogy: this essay seems to suggest that within the social media wellbeing research community, it is a common practice to bundle all interactions with social media into a catch-all “use” or “time spent” quantity. I’m not sure this was ever a great methodological choice, but it seems even less useful to do that when studying the effects of AI, because the set of things one can do with AI is much broader, even compared to social media. Imagine 2 study subjects: one who uses AI to speed through school assignments, versus another who uses AI to create practice problems to then solve independently. Under a coarse-grained “use” metric, these two subjects could very well be classified into the same bucket, which feels obviously wrong. There are inertial forces (e.g. the fact that it’s relatively easy to get self-reports of time spent on a specific AI product) that will likely push researchers toward the catch-all “use” value as a predictor variable, but ultimately that practice does not seem like it should be rolled over. To me, this case study showcases the limits of analogical thinking about how past technologies have shaped human cognition and behavior. Ultimately, I actually found it a more useful exercise to think about the ways in which AI is not like social media, than to try to draw explicit parallels. For a more well-formed example of what this contrastive approach can buy you: we can have more secure foundations to reason about how LLMs could reshape public opinion, by contrasting it with the effects of social media. If we hope to arrive at a holistic picture of how AI is changing how we think, there are a lot of hard problems that need to be addressed; social media research didn’t pave the way methodologically, although maybe it offers a useful point of contrast. An additional consolation is that the object of study is so flexible as to be a useful analytical tool to help us answer these questions, more quickly: AI, used judiciously, could potentially help unblock some of these methodological challenges (e.g. coding agents could make it easy to build a more granular tracker of individuals’ interactions with AI). The big question remains, though, of how we will sift through an ever-expanding body of AI-powered research to understand what this technology is doing to our own minds - and if we can draw conclusions at the speed at which AI is infusing itself into our lives.]]></summary></entry><entry><title type="html">Escaping the verifiability trap in AI simulations of human behavior</title><link href="https://dharakyu.com/blog/2026/04/27/escaping-the-verifiability-trap/" rel="alternate" type="text/html" title="Escaping the verifiability trap in AI simulations of human behavior" /><published>2026-04-27T00:00:00+00:00</published><updated>2026-04-27T00:00:00+00:00</updated><id>https://dharakyu.com/blog/2026/04/27/escaping-the-verifiability-trap</id><content type="html" xml:base="https://dharakyu.com/blog/2026/04/27/escaping-the-verifiability-trap/"><![CDATA[<p>Can you study human behavior without collecting human data? Just a few years ago, this would have seemed like a completely nonsensical proposition. But it has become top-of-mind for me (and many others), given growing interest in the use of LLM-based approaches for simulating human behavior in surveys and lab tasks. This is not just an abstract academic exercise: many startups are claiming to be able to simulate customer behavior in place of conducting traditional user research, and there is widespread interest among model developers in building <a href="https://jessylin.com/2025/09/25/user-simulators-2/">better user simulators</a> for LLM post-training.</p>

<p>Reflecting all this interest, there is now an emerging field studying the suitability of LLM agents as drop-in replacements for human participants in social science experiments. Most empirical papers take the following form. First, find an existing experiment (or set of experiments) that was conducted with real human participants. Then, re-run the experiment by presenting those experimental stimuli to an LLM agent and collecting its responses (possibly conditioning on a “persona” profile with demographic information to elicit these responses). Finally, measure the overlap between the simulated behavior and the ground-truth human participant data - an approach referred to <a href="https://arxiv.org/abs/2602.15785"><em>heuristic validation</em></a> by Hullman et al.</p>

<p>This research program of measuring agreement between model simulation and human behavior inherently selects for the use of experiments where it is straightforward to measure this agreement. In the <a href="https://www.nature.com/articles/s41586-025-09215-4">Centaur paper</a>, which measured alignment between LLM predictions and human experimental data across a variety of psychological tasks, the most-represented experimental task (comprising over 40% of all human behavioral trial outcomes used to fine-tune the model) is intertemporal choice. In this task, participants select between receiving some amount of money now, versus a (greater) amount at a later time - a decision problem with 2 fixed options. Similarly, <a href="https://docsend.com/view/qeeccuggec56k9hd">in another study</a> (still under review as of this writing), the primary set of experiments used to evaluate LLM fidelity were surveys in which participants indicated their responses via a numerical rating scale (e.g. a 7-point scale from “strongly disagree” to “strongly agree”). An LLM can be straightforwardly prompted to select A or B or produce a number between 1 and 7, producing outputs that are immediately comparable to the original empirical data. To use the parlance of AI research, there is a structural bias for <em>verifiable</em> domains - the types of experiments models are best suited to simulate are the ones where the faithfulness of that simulation (e.g., the “correctness” of the model’s generation) can be readily computed with off-the-shelf metrics.</p>

<p>Given the boldness of the idea being evaluated - literally removing humans from human behavioral experiments - this gravitation toward close-ended tasks seems eminently reasonable; we should be able to evaluate these claims quantitatively. But there is one core problem: the types of experiments that are most straightforward to simulate are the most constrained and simplified ones, which often bear the most tenuous connections to the cognitive capabilities and behaviors we presumably would want to capture in simulation.</p>

<h2 id="an-old-problem-resurfaced">An old problem, resurfaced</h2>

<p>A structural challenge in psychological science is the fact that theoretical constructs of interest cannot directly be measured: instead, an experimenter must come up with a way to operationalize the construct in a way that <em>can</em> be instrumented. For example, if I am interested in <em>social reasoning</em>, then I might operationalize this by conducting an experiment to evaluate the types of inferences people make about other people’s mental states in a variety of scenarios.</p>

<p>Within cognitive psychology, this process of operationalization has traditionally involved the use of experimental tasks which serve as testbeds for studying the phenomenon of interest. One common thread - and key benefit - of these tasks is that they tend to constrain the set of possible behaviors in a way that makes it easy to verify the predictive accuracy of a statistical or cognitive model.</p>

<p>The core downside, however, is that these tasks often abstract away parts of the problem that could be consequential for understanding the original motivating phenomenon. To see this, let’s revisit the example of intertemporal choice. There are many real-world decisions that are problems of intertemporal choice, like whether one should spend a paycheck or save for the future, or go out with friends or prepare for an important meeting. But the experimental proxy setup - picking between a set of fixed options, each of which has an unambiguous value in the context of the study - is a highly stylized representation. In the real-world setting, people must integrate complex information and internally construct value signals (e.g. the value of catching up with friends versus the value of impressing your boss); these processes presumably also affect how people arrive at their decisions. This is <em>not</em> to say that these kinds of tasks are not useful; some degree of abstraction is inevitable and can often be highly instructive. But it raises <a href="https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/article/generalizability-crisis/AD386115BA539A759ACB3093760F4824">questions about the generalizability of such studies</a> - what inferences are licensed about underlying cognitive mechanisms, given the gap between the experimental setting and the real-world behavior?</p>

<p>Back to the question of simulating human behavior: these sorts of experiments, with the most stylized operationalizations of human behavior, are disproportionately the ones being used to adjudicate the claim that LLMs can or cannot simulate human behavior. And so these same questions about generalizability arise with simulation: if a model can predict human responses in a simplified laboratory experiment of 2-way decision-making, what can we conclude about that model’s capability to simulate human behavior across other classes of decisions? Without appropriate scoping, we might make unjustified <a href="https://arxiv.org/abs/2602.15785">inferential leaps</a> and ultimately <a href="https://www.sciencedirect.com/science/article/abs/pii/S1364661325002517">mislead ourselves</a> about the extent to which LLMs can capture human behavior.</p>

<p>The problem isn’t LLMs, per se: rather, their proliferation is exposing deeper issues of ecological validity in social science experiments. Experimental psychologists have faced the trap of verifiable domains long before that term emerged. But this time around, the stakes are arguably higher, because methods and paradigms from the field are being used to make broader claims about a technology with a reach far beyond academic science.</p>

<p>There could, however, be a way out of this morass: the same tools that are surfacing this problem are the ones that could help us develop solutions, by giving us qualitatively new ways to analyze naturalistic human behavior at scale.</p>

<h2 id="closing-the-gap-between-the-task-and-the-phenomenon">Closing the gap between the task and the phenomenon</h2>

<p>If the goal is to ultimately predict “real world” behavior, a clear next step is to study human behavior in settings that more closely reflect the dynamics of the original motivating phenomena. In the past, this was an extremely difficult thing to ask of scientists, because there was an inherent tradeoff between naturalism and quantitative analysis: you could study people doing the most complex thing in the world, but you wouldn’t have (quantitative) tools to make sense of those behaviors.</p>

<p>LLMs (<a href="https://arxiv.org/abs/2502.20349">and AI more broadly</a>) fundamentally change that calculus: it is now possible to extract much richer dependent variables from more naturalistic and open-ended traces of behavior, therefore expanding the complexity and scope of our (human) experimental paradigms. To provide an example: in some recent work of my own, we collected a large dataset of free-form natural language conversations among pairs of human participants tasked with solving an open-ended joint planning problem; these pairs then were tasked with implementing their joint plan. We used an LLM to convert the structured agreements participants expressed in language into Python program representations. These program-like representations allowed us to simulate how a pair with a given strategy might behave on a different set of inputs. This in turn made it possible to directly compare strategy use across pairs and therefore to track how the internal algorithmic structure of the agreements changed over time - a dependent variable that simply would have not been feasible to measure in the pre-LLM era.</p>

<p>Of course, this is just one study, and it does not solve all the methodological challenges I previously outlined: this task is not a fully comprehensive proxy for how people engage in joint planning in the “real world”. But it helps, I hope, to burnish the case that we no longer have to choose between greater naturalism in experimental paradigms, and quantitative analysis: we increasingly can have both.</p>

<h2 id="simulation-beyond-verifiable-domains">Simulation beyond verifiable domains</h2>

<p>Once we build tools to characterize more complex, in-the-wild human behaviors, we will be better equipped to revisit the simulation debate and evaluate if LLMs can faithfully capture said behaviors. When conducting human experiments, we are often left wondering if theories that describe the pattern of findings in a particular experiment have any explanatory power beyond the lab setup. With LLMs, we can actually evaluate this quantitatively: we can test if models that are optimized to recapitulate behavior in laboratory tasks can also predict behavior in more naturalistic settings.</p>

<p>This is a key benefit afforded by LLMs, compared to previous classes of cognitive models; they can produce much more open-ended outputs (i.e. anything that can be expressed in natural language). So the same model that is optimized to match the distribution of behavior in binary choice lab experiments (through fine-tuning, prompting, scaffolding within a larger system, or any combination of these interventions) can also be used to simulate how people make more open-ended decisions. <a href="https://karpathy.bearblog.dev/verifiability/">Given everything we’ve observed about AI progress</a>, it seems within reason that a model could eventually predict with a high degree of accuracy human decisions in forced-choice laboratory experiments: the more interesting question is if this same model could capture behavior in other classes of decision problems.</p>

<p>Of course, this also raises the question of where one could source these more “naturalistic” datasets. One possibility is to run more complex experiments that elicit more open-ended forms of behavior, as mentioned earlier. Although there would still be concerns about the artificiality of an experimenter-prompted lab experiment, the actual tasks themselves could retain more of the character of the original motivating problems. Another possible trove is conversational data from social media, as tech platforms have incidentally become a <a href="https://cocosci.princeton.edu/tom/papers/ComputationalCognitiveRevolution.pdf">rich repository of human behavioral data</a>. Given the broadness of activities on social media, using this kind of data necessitates filtering to select subsets that are relevant for the question at hand. One example of this comes from Sudeep Bhatia et al., who <a href="https://www.pnas.org/doi/full/10.1073/pnas.2406489122">developed a pipeline</a> to extract a corpus of personal life decisions from a Reddit corpus. They then used LLMs to extract feature representations for these decision problems, to test if existing theories of decision-making could explain the choices observed in the Reddit dataset.</p>

<p>If we’re dealing with more open-ended forms of data and can’t use traditional metrics, how do we compare open-ended LLM simulations to ground-truth human behavioral data? LLMs can again be part of the solution: the same pipeline for characterizing structure in human behavior can be repurposed to characterize structure in simulator outputs, allowing for direct comparisons. Here, we’re drawing a distinction between <em>verifiable</em> and <em>measurable</em>: while measuring agreement between open-ended human behavior and agent simulations is not generally something that can be reduced to a single standardized metric, it is still possible to instrument it with rigor and quantitative precision. Of course, this still requires careful craftsmanship on the part of the (human) scientist, e.g. by identifying a suitable dependent variable, clearly operationalizing a definition of similarity with respect to that dependent variable, providing high-quality example annotations, and validating that LLM judgements actually align with experimenter-defined criteria.</p>

<p>In a recent essay, Hiranya Peiris <a href="https://www.nature.com/articles/s41550-026-02837-2">made the case</a> that the advent of LLMs in scientific practice is not actually creating qualitatively new problems, but rather exposing longstanding ones. Issues such as the explosion of low quality papers and the scientific shallowness of “data science for x” approaches have existed long before LLMs burst onto the scene. A similar dynamic is emerging in the discourse around LLM for simulating human behavior: we are rediscovering core methodological problems that social science has been grappling with for decades.</p>

<p>There is one (bleak) future in which the use of LLMs for simulation leads us to “<a href="https://www.nature.com/articles/d41586-025-01067-2">squeeze more predictive juice out of flawed theories and inadequate paradigms</a>”, as put vividly by Sayash Kapoor and Arvind Narayanan. But I think there is reason for optimism, because these tools fundamentally extend what is measurable and therefore expand the types of questions we can ask about complex behavioral datasets. AI could fulfill its promise of transforming social science - if viewed not merely as a way to scale up what’s been done, but rather to measure what we never could before.</p>

<hr />

<p><em>Thank you to <a href="https://www.normanmu.com/">Norman</a>, <a href="https://www.ivangrahek.com/">Ivan</a>, Karthik and Mathew for helpful feedback and discussion!</em></p>]]></content><author><name></name></author><summary type="html"><![CDATA[Can you study human behavior without collecting human data? Just a few years ago, this would have seemed like a completely nonsensical proposition. But it has become top-of-mind for me (and many others), given growing interest in the use of LLM-based approaches for simulating human behavior in surveys and lab tasks. This is not just an abstract academic exercise: many startups are claiming to be able to simulate customer behavior in place of conducting traditional user research, and there is widespread interest among model developers in building better user simulators for LLM post-training. Reflecting all this interest, there is now an emerging field studying the suitability of LLM agents as drop-in replacements for human participants in social science experiments. Most empirical papers take the following form. First, find an existing experiment (or set of experiments) that was conducted with real human participants. Then, re-run the experiment by presenting those experimental stimuli to an LLM agent and collecting its responses (possibly conditioning on a “persona” profile with demographic information to elicit these responses). Finally, measure the overlap between the simulated behavior and the ground-truth human participant data - an approach referred to heuristic validation by Hullman et al. This research program of measuring agreement between model simulation and human behavior inherently selects for the use of experiments where it is straightforward to measure this agreement. In the Centaur paper, which measured alignment between LLM predictions and human experimental data across a variety of psychological tasks, the most-represented experimental task (comprising over 40% of all human behavioral trial outcomes used to fine-tune the model) is intertemporal choice. In this task, participants select between receiving some amount of money now, versus a (greater) amount at a later time - a decision problem with 2 fixed options. Similarly, in another study (still under review as of this writing), the primary set of experiments used to evaluate LLM fidelity were surveys in which participants indicated their responses via a numerical rating scale (e.g. a 7-point scale from “strongly disagree” to “strongly agree”). An LLM can be straightforwardly prompted to select A or B or produce a number between 1 and 7, producing outputs that are immediately comparable to the original empirical data. To use the parlance of AI research, there is a structural bias for verifiable domains - the types of experiments models are best suited to simulate are the ones where the faithfulness of that simulation (e.g., the “correctness” of the model’s generation) can be readily computed with off-the-shelf metrics. Given the boldness of the idea being evaluated - literally removing humans from human behavioral experiments - this gravitation toward close-ended tasks seems eminently reasonable; we should be able to evaluate these claims quantitatively. But there is one core problem: the types of experiments that are most straightforward to simulate are the most constrained and simplified ones, which often bear the most tenuous connections to the cognitive capabilities and behaviors we presumably would want to capture in simulation. An old problem, resurfaced A structural challenge in psychological science is the fact that theoretical constructs of interest cannot directly be measured: instead, an experimenter must come up with a way to operationalize the construct in a way that can be instrumented. For example, if I am interested in social reasoning, then I might operationalize this by conducting an experiment to evaluate the types of inferences people make about other people’s mental states in a variety of scenarios. Within cognitive psychology, this process of operationalization has traditionally involved the use of experimental tasks which serve as testbeds for studying the phenomenon of interest. One common thread - and key benefit - of these tasks is that they tend to constrain the set of possible behaviors in a way that makes it easy to verify the predictive accuracy of a statistical or cognitive model. The core downside, however, is that these tasks often abstract away parts of the problem that could be consequential for understanding the original motivating phenomenon. To see this, let’s revisit the example of intertemporal choice. There are many real-world decisions that are problems of intertemporal choice, like whether one should spend a paycheck or save for the future, or go out with friends or prepare for an important meeting. But the experimental proxy setup - picking between a set of fixed options, each of which has an unambiguous value in the context of the study - is a highly stylized representation. In the real-world setting, people must integrate complex information and internally construct value signals (e.g. the value of catching up with friends versus the value of impressing your boss); these processes presumably also affect how people arrive at their decisions. This is not to say that these kinds of tasks are not useful; some degree of abstraction is inevitable and can often be highly instructive. But it raises questions about the generalizability of such studies - what inferences are licensed about underlying cognitive mechanisms, given the gap between the experimental setting and the real-world behavior? Back to the question of simulating human behavior: these sorts of experiments, with the most stylized operationalizations of human behavior, are disproportionately the ones being used to adjudicate the claim that LLMs can or cannot simulate human behavior. And so these same questions about generalizability arise with simulation: if a model can predict human responses in a simplified laboratory experiment of 2-way decision-making, what can we conclude about that model’s capability to simulate human behavior across other classes of decisions? Without appropriate scoping, we might make unjustified inferential leaps and ultimately mislead ourselves about the extent to which LLMs can capture human behavior. The problem isn’t LLMs, per se: rather, their proliferation is exposing deeper issues of ecological validity in social science experiments. Experimental psychologists have faced the trap of verifiable domains long before that term emerged. But this time around, the stakes are arguably higher, because methods and paradigms from the field are being used to make broader claims about a technology with a reach far beyond academic science. There could, however, be a way out of this morass: the same tools that are surfacing this problem are the ones that could help us develop solutions, by giving us qualitatively new ways to analyze naturalistic human behavior at scale. Closing the gap between the task and the phenomenon If the goal is to ultimately predict “real world” behavior, a clear next step is to study human behavior in settings that more closely reflect the dynamics of the original motivating phenomena. In the past, this was an extremely difficult thing to ask of scientists, because there was an inherent tradeoff between naturalism and quantitative analysis: you could study people doing the most complex thing in the world, but you wouldn’t have (quantitative) tools to make sense of those behaviors. LLMs (and AI more broadly) fundamentally change that calculus: it is now possible to extract much richer dependent variables from more naturalistic and open-ended traces of behavior, therefore expanding the complexity and scope of our (human) experimental paradigms. To provide an example: in some recent work of my own, we collected a large dataset of free-form natural language conversations among pairs of human participants tasked with solving an open-ended joint planning problem; these pairs then were tasked with implementing their joint plan. We used an LLM to convert the structured agreements participants expressed in language into Python program representations. These program-like representations allowed us to simulate how a pair with a given strategy might behave on a different set of inputs. This in turn made it possible to directly compare strategy use across pairs and therefore to track how the internal algorithmic structure of the agreements changed over time - a dependent variable that simply would have not been feasible to measure in the pre-LLM era. Of course, this is just one study, and it does not solve all the methodological challenges I previously outlined: this task is not a fully comprehensive proxy for how people engage in joint planning in the “real world”. But it helps, I hope, to burnish the case that we no longer have to choose between greater naturalism in experimental paradigms, and quantitative analysis: we increasingly can have both. Simulation beyond verifiable domains Once we build tools to characterize more complex, in-the-wild human behaviors, we will be better equipped to revisit the simulation debate and evaluate if LLMs can faithfully capture said behaviors. When conducting human experiments, we are often left wondering if theories that describe the pattern of findings in a particular experiment have any explanatory power beyond the lab setup. With LLMs, we can actually evaluate this quantitatively: we can test if models that are optimized to recapitulate behavior in laboratory tasks can also predict behavior in more naturalistic settings. This is a key benefit afforded by LLMs, compared to previous classes of cognitive models; they can produce much more open-ended outputs (i.e. anything that can be expressed in natural language). So the same model that is optimized to match the distribution of behavior in binary choice lab experiments (through fine-tuning, prompting, scaffolding within a larger system, or any combination of these interventions) can also be used to simulate how people make more open-ended decisions. Given everything we’ve observed about AI progress, it seems within reason that a model could eventually predict with a high degree of accuracy human decisions in forced-choice laboratory experiments: the more interesting question is if this same model could capture behavior in other classes of decision problems. Of course, this also raises the question of where one could source these more “naturalistic” datasets. One possibility is to run more complex experiments that elicit more open-ended forms of behavior, as mentioned earlier. Although there would still be concerns about the artificiality of an experimenter-prompted lab experiment, the actual tasks themselves could retain more of the character of the original motivating problems. Another possible trove is conversational data from social media, as tech platforms have incidentally become a rich repository of human behavioral data. Given the broadness of activities on social media, using this kind of data necessitates filtering to select subsets that are relevant for the question at hand. One example of this comes from Sudeep Bhatia et al., who developed a pipeline to extract a corpus of personal life decisions from a Reddit corpus. They then used LLMs to extract feature representations for these decision problems, to test if existing theories of decision-making could explain the choices observed in the Reddit dataset. If we’re dealing with more open-ended forms of data and can’t use traditional metrics, how do we compare open-ended LLM simulations to ground-truth human behavioral data? LLMs can again be part of the solution: the same pipeline for characterizing structure in human behavior can be repurposed to characterize structure in simulator outputs, allowing for direct comparisons. Here, we’re drawing a distinction between verifiable and measurable: while measuring agreement between open-ended human behavior and agent simulations is not generally something that can be reduced to a single standardized metric, it is still possible to instrument it with rigor and quantitative precision. Of course, this still requires careful craftsmanship on the part of the (human) scientist, e.g. by identifying a suitable dependent variable, clearly operationalizing a definition of similarity with respect to that dependent variable, providing high-quality example annotations, and validating that LLM judgements actually align with experimenter-defined criteria. In a recent essay, Hiranya Peiris made the case that the advent of LLMs in scientific practice is not actually creating qualitatively new problems, but rather exposing longstanding ones. Issues such as the explosion of low quality papers and the scientific shallowness of “data science for x” approaches have existed long before LLMs burst onto the scene. A similar dynamic is emerging in the discourse around LLM for simulating human behavior: we are rediscovering core methodological problems that social science has been grappling with for decades. There is one (bleak) future in which the use of LLMs for simulation leads us to “squeeze more predictive juice out of flawed theories and inadequate paradigms”, as put vividly by Sayash Kapoor and Arvind Narayanan. But I think there is reason for optimism, because these tools fundamentally extend what is measurable and therefore expand the types of questions we can ask about complex behavioral datasets. AI could fulfill its promise of transforming social science - if viewed not merely as a way to scale up what’s been done, but rather to measure what we never could before. Thank you to Norman, Ivan, Karthik and Mathew for helpful feedback and discussion!]]></summary></entry><entry><title type="html">Building the boat while sailing: studying social interaction in 2026</title><link href="https://dharakyu.com/blog/2026/03/12/building-the-boat-while-sailing/" rel="alternate" type="text/html" title="Building the boat while sailing: studying social interaction in 2026" /><published>2026-03-12T00:00:00+00:00</published><updated>2026-03-12T00:00:00+00:00</updated><id>https://dharakyu.com/blog/2026/03/12/building-the-boat-while-sailing</id><content type="html" xml:base="https://dharakyu.com/blog/2026/03/12/building-the-boat-while-sailing/"><![CDATA[<p>I have - slowly but surely - been making progress toward a long-standing goal of mine: reading Tolstoy’s <em>War and Peace</em>. It’s far from the slog I remember when I last tried to read it (~10 years ago). Instead I find myself drawn in by the evocative portraits drawn of the novel’s central characters, in particular the ornate descriptions of their beliefs and desires - Prince Andrei’s stoicism, Pierre’s search for purpose, Natasha’s romantic fervor - which color how they perceive the world.</p>

<p>This degree of interiority is striking. Rarely, it seems, do we have such unfettered access to inner machinations and traces of reasoning situated in other minds. But even if we don’t have Tolstoy on our shoulder to offer astute insights into our day-to-day interactions, we do have something else: the ability to infer otherwise inaccessible mental states through interactive conversation.</p>

<p>To see this, imagine helping a friend with a research talk. A key part of the presentation initially comes across as unclear, but after an extended back-and-forth with carefully targeted questions and clarifications, you both converge on the fleshed-out form of the idea. Crucially, this resolution can only be reached after multiple exchanges; it’s not a one-shot thing. More broadly, this example illustrates that while we cannot literally read each others’ minds, the affordance of language gives us a way to express thoughts, and a mechanism to coax them out through unfolding interaction.</p>

<p>For the past few years, I’ve been thinking about the cognitive principles that enable these dynamic forms of social interaction in intelligent agents: how is it that people can accomplish feats of joint action unparalleled by any other animal species, things like building a skyscraper, playing in a band or establishing the United Nations? Increasingly, I’ve also been thinking about how new, distinct forms of intelligence - namely, AI agents built on top of large language models - interact with people and with each other, and what might result from these encounters.</p>

<p>The obvious common thread across all these classes of interactions is the use of language to convey meaning and to shape other’s behavior. Yet for all the clear importance of interactive language use - both as a window into what people are thinking, and as an action-oriented tool for accomplishing goals - cognitive science is still very much in the early stages of study.</p>

<p>In the rest of this essay, I will lay out some thoughts about possible future directions for the science of interaction, broadly construed. My hope is that these directions can help us (in a small way) answer long-standing questions about the nature of human cognition, and provide some methodological grounding to study what is going on in our interactions with these new forms of intelligence.</p>

<h2 id="the-cognitive-science-of-human-social-interaction">The cognitive science of (human) social interaction</h2>

<p>Human social interaction is largely realized through language; it’s the medium that we use to convey everything from immediate practical concerns to abstract ideas that transcend space and time. The corollary is that interaction is often the impetus for language use. If you perform a task only for yourself, there’s less of a reason to create an externalized trace of that in language, either spoken or written (although there are notable exceptions, such as journal entries in a diary). Accordingly, much of the language that we are exposed to reflects this explicitly multi-agent purpose of transmitting information from one mind to another, or jointly constructing new knowledge.</p>

<p>There are academic fields devoted to the study of strategic interaction sans language (e.g., game theory). Within cognitive and psychological science, researchers studying nonlinguistic populations (very young children and nonhuman primates) have used techniques such as gaze analysis to study interactions. For full-fledged natural language conversational data, important methodological foundations have been established, perhaps most notably by practitioners of <a href="https://methods.sagepub.com/book/mono/preview/conversation-analysis.pdf">conversational analysis</a>. But the surface has barely been scratched, even though language is arguably the richest behavioral data we can collect, the most high-resolution readout of otherwise inaccessible mental processes.</p>

<p>In fact, <a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC10932833/">cognitive and psychological science has not really focused on the study of interaction, period</a>. Even subareas of cognitive science studying social cognition <a href="https://www.annualreviews.org/content/journals/10.1146/annurev-devpsych-111323-115032">have traditionally studied human sociality in a kind of asocial way</a>, e.g. making inferences about the mental states of a single other person. Social psychology faces a similar dynamic, in that it’s often focused on documenting internal attitudes and beliefs; the unit of analysis is often not the interaction itself.</p>

<p>In this sense, psychology is the least social of the social sciences. Although this feels like an oversight, it’s hard to blame past researchers because historically, it has been methodologically very challenging to study interaction. Experiments in groups are more difficult to run than experiments with individuals, and the ensuing traces of interaction in language were hard to analyze at scale.</p>

<p>This latter challenge, of course, has now changed with LLMs. It is now doable to analyze open-ended language in a very fine-grained way: language models <a href="https://www.pnas.org/doi/abs/10.1073/pnas.2308950121">can be used to identify specific attributes in language data</a> with a high degree of precision, or transform language into other formal representations more suitable for quantitative analysis. There is a lot of exciting work on <a href="https://arxiv.org/abs/2510.26202">new LLM-powered methods to extract structured signal</a> in text data in principled ways.</p>

<p>But there’s a strange sort of circularity: the very tools that facilitate the study of interaction are becoming objects of study in their own right. LLMs are interacting with human users through various chatbot products, and increasingly in agentic setups with people and other agentic systems. Tasks that have been traditionally thought of as solitary endeavors (e.g. programming) are becoming fundamentally interactive; modern software engineering now entails iterating on a spec through back-and-forths with agents.</p>

<p>I am far from the first person to point out that AI models can occupy multiple distinct roles in behavioral science, both as a tool and a model system (<a href="https://static1.squarespace.com/static/53d29678e4b04e06965e9423/t/6566e06ac95b0b61f8810a99/1701240942374/2023+--+LLMs+psychology.pdf">some</a> <a href="https://www.nature.com/articles/s41586-024-07146-0">examples</a> <a href="https://arxiv.org/abs/2506.00052">here</a> and <a href="https://www.annualreviews.org/content/journals/10.1146/annurev-psych-030625-040748">here</a>). But what I would argue is different about this particular phenomenon is that the scientific study of social interaction is uniquely inchoate, because of the historical blockers to studying it. We’re building the boat while sailing: we’re trying to develop tools to study these kinds of interactions as they are already changing our workflows and ultimately shaping how we think.</p>

<h2 id="understanding-new-types-of-interactions">Understanding new types of interactions</h2>

<p>The rise of interactive AI systems has revealed gaps in our understanding of interaction, broadly construed, and has created new challenges of shared interest across many fields. I see this as a call-to-arms and opportunity for cognitive science. This presents a chance to develop new tools, frameworks, models and methods that not only help resolve longstanding theoretical questions about representation and computation, but can also be brought to bear on questions of broad societal relevance.</p>

<p>Just as the development of information technology and feedback systems during World War II exposed new fundamental questions about computation and led to scientific breakthroughs in information theory and cybernetics (control theory), today’s conversational AI systems can play a catalyzing role for the science of interactive intelligence, which transcends traditional disciplinary boundaries. In an ideal world, we’d work toward a science of interaction that offers a sufficiently broad and rigorous toolkit that is useful across all classes of interactions - purely human, <a href="https://theaidigest.org/village">purely AI</a> and human-AI teams - and responsive to the ways in which these interactions will change, as model capabilities improve and human behavior changes.</p>

<p>As a starting point, it may be useful to reflect on what’s shared and what’s different across these classes of interactions. Of course, the involved agents themselves - and the process by which they acquire worldly knowledge - can be fundamentally different. Many have pointed out that <a href="https://www.sciencedirect.com/science/article/pii/S1364661323002036">the scale and content of training data differs across people and AI models</a>. Yet there are high-level structural similarities. Analogous to how a person might listen to conversation between others, LLMs in pre-training learn from traces of interactive reasoning, where the learning agent is not an involved party in the interaction. Even before all the post-training explicitly designed to make LLMs more useful as “interactive” conversational agents, they are exposed to interactions between multiple individuals with distinct epistemic states exchanging information over an unfolding conversation, played out in comment threads and online forums. In this sense, both humans and AI models occupy the role of an observer embedded in a society of interacting agents, uptaking the externalized traces.</p>

<p>The form factor of the interaction itself is also worth considering. As noted earlier, all classes of interactions are primarily realized through natural language - a point of commonality. Yet there are also clear differences, e.g. in the availability and latency of responses - a chatbot can answer instantaneously, without incurring a “cost” (in a psycholinguistic, not necessarily monetary, sense) to produce such a response, becoming a sort of assistant on demand.</p>

<p>Given this reduction in barriers, one of the most promising, although also potentially perilous, opportunities for interactive LLM use is as a tool for personalized learning. For most of human history, before widespread literacy, <a href="https://www.derekthompson.org/p/why-the-decline-of-literacyand-the">learning through the experiences of others was fundamentally interactive</a>: there was no way to learn from another person other than being physically co-located and speaking directly with them. <a href="https://www.theatlantic.com/ideas/archive/2025/10/ai-deskilling-automation-technology/684669/">The advent of the written word changed that dynamic</a>; all kinds of knowledge, from mathematics to poisonous mushroom species, became accessible in written, static texts that were broadly accessible. Learning through AI interaction feels like a kind of interpolation between these different modes, in that it’s primarily realized through the familiar form of written text, yet the text is no longer static; it is generated through unfolding interaction with a teacher-like figure. In this sense, learning with LLMs is both an extension of existing paradigms and a qualitatively new phenomenon.</p>

<p>Of course, these are very high-level reflections on the structural similarities and differences across different classes of interactions - what would it mean to actually interrogate this rigorously? Forging this path will likely involve revisiting broader and perpetually-unresolved methodological debates on <a href="https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/article/beyond-playing-20-questions-with-nature-integrative-experiment-design-in-the-social-and-behavioral-sciences/7E0D34D5AE2EFB9C0902414C23E0C292">experimental control</a> vs. <a href="https://infinitefaculty.substack.com/p/what-cognitive-science-can-learn">naturalism</a>, and formal modeling vs. qualitative insight - but one thing that seems clear is the need to design experiments that expose open-ended interaction in the first place, or otherwise aggregate naturally-occurring datasets that reveal traces of interaction. Maybe this seems like kind of a trivial point, but at the same time, <a href="https://onlinelibrary.wiley.com/doi/full/10.1111/cogs.13230">much remains to be done in investigating these phenomena through the lens of cognitive science</a>. It feels more attainable—and more important—than ever to study this fundamental aspect of intelligence.</p>

<hr />

<p><em>Thank you to <a href="https://billdthompson.github.io/">Bill Thompson</a> for feedback on earlier drafts, as well as to the many people who have helped me develop the ideas here!</em></p>]]></content><author><name></name></author><summary type="html"><![CDATA[I have - slowly but surely - been making progress toward a long-standing goal of mine: reading Tolstoy’s War and Peace. It’s far from the slog I remember when I last tried to read it (~10 years ago). Instead I find myself drawn in by the evocative portraits drawn of the novel’s central characters, in particular the ornate descriptions of their beliefs and desires - Prince Andrei’s stoicism, Pierre’s search for purpose, Natasha’s romantic fervor - which color how they perceive the world. This degree of interiority is striking. Rarely, it seems, do we have such unfettered access to inner machinations and traces of reasoning situated in other minds. But even if we don’t have Tolstoy on our shoulder to offer astute insights into our day-to-day interactions, we do have something else: the ability to infer otherwise inaccessible mental states through interactive conversation. To see this, imagine helping a friend with a research talk. A key part of the presentation initially comes across as unclear, but after an extended back-and-forth with carefully targeted questions and clarifications, you both converge on the fleshed-out form of the idea. Crucially, this resolution can only be reached after multiple exchanges; it’s not a one-shot thing. More broadly, this example illustrates that while we cannot literally read each others’ minds, the affordance of language gives us a way to express thoughts, and a mechanism to coax them out through unfolding interaction. For the past few years, I’ve been thinking about the cognitive principles that enable these dynamic forms of social interaction in intelligent agents: how is it that people can accomplish feats of joint action unparalleled by any other animal species, things like building a skyscraper, playing in a band or establishing the United Nations? Increasingly, I’ve also been thinking about how new, distinct forms of intelligence - namely, AI agents built on top of large language models - interact with people and with each other, and what might result from these encounters. The obvious common thread across all these classes of interactions is the use of language to convey meaning and to shape other’s behavior. Yet for all the clear importance of interactive language use - both as a window into what people are thinking, and as an action-oriented tool for accomplishing goals - cognitive science is still very much in the early stages of study. In the rest of this essay, I will lay out some thoughts about possible future directions for the science of interaction, broadly construed. My hope is that these directions can help us (in a small way) answer long-standing questions about the nature of human cognition, and provide some methodological grounding to study what is going on in our interactions with these new forms of intelligence. The cognitive science of (human) social interaction Human social interaction is largely realized through language; it’s the medium that we use to convey everything from immediate practical concerns to abstract ideas that transcend space and time. The corollary is that interaction is often the impetus for language use. If you perform a task only for yourself, there’s less of a reason to create an externalized trace of that in language, either spoken or written (although there are notable exceptions, such as journal entries in a diary). Accordingly, much of the language that we are exposed to reflects this explicitly multi-agent purpose of transmitting information from one mind to another, or jointly constructing new knowledge. There are academic fields devoted to the study of strategic interaction sans language (e.g., game theory). Within cognitive and psychological science, researchers studying nonlinguistic populations (very young children and nonhuman primates) have used techniques such as gaze analysis to study interactions. For full-fledged natural language conversational data, important methodological foundations have been established, perhaps most notably by practitioners of conversational analysis. But the surface has barely been scratched, even though language is arguably the richest behavioral data we can collect, the most high-resolution readout of otherwise inaccessible mental processes. In fact, cognitive and psychological science has not really focused on the study of interaction, period. Even subareas of cognitive science studying social cognition have traditionally studied human sociality in a kind of asocial way, e.g. making inferences about the mental states of a single other person. Social psychology faces a similar dynamic, in that it’s often focused on documenting internal attitudes and beliefs; the unit of analysis is often not the interaction itself. In this sense, psychology is the least social of the social sciences. Although this feels like an oversight, it’s hard to blame past researchers because historically, it has been methodologically very challenging to study interaction. Experiments in groups are more difficult to run than experiments with individuals, and the ensuing traces of interaction in language were hard to analyze at scale. This latter challenge, of course, has now changed with LLMs. It is now doable to analyze open-ended language in a very fine-grained way: language models can be used to identify specific attributes in language data with a high degree of precision, or transform language into other formal representations more suitable for quantitative analysis. There is a lot of exciting work on new LLM-powered methods to extract structured signal in text data in principled ways. But there’s a strange sort of circularity: the very tools that facilitate the study of interaction are becoming objects of study in their own right. LLMs are interacting with human users through various chatbot products, and increasingly in agentic setups with people and other agentic systems. Tasks that have been traditionally thought of as solitary endeavors (e.g. programming) are becoming fundamentally interactive; modern software engineering now entails iterating on a spec through back-and-forths with agents. I am far from the first person to point out that AI models can occupy multiple distinct roles in behavioral science, both as a tool and a model system (some examples here and here). But what I would argue is different about this particular phenomenon is that the scientific study of social interaction is uniquely inchoate, because of the historical blockers to studying it. We’re building the boat while sailing: we’re trying to develop tools to study these kinds of interactions as they are already changing our workflows and ultimately shaping how we think. Understanding new types of interactions The rise of interactive AI systems has revealed gaps in our understanding of interaction, broadly construed, and has created new challenges of shared interest across many fields. I see this as a call-to-arms and opportunity for cognitive science. This presents a chance to develop new tools, frameworks, models and methods that not only help resolve longstanding theoretical questions about representation and computation, but can also be brought to bear on questions of broad societal relevance. Just as the development of information technology and feedback systems during World War II exposed new fundamental questions about computation and led to scientific breakthroughs in information theory and cybernetics (control theory), today’s conversational AI systems can play a catalyzing role for the science of interactive intelligence, which transcends traditional disciplinary boundaries. In an ideal world, we’d work toward a science of interaction that offers a sufficiently broad and rigorous toolkit that is useful across all classes of interactions - purely human, purely AI and human-AI teams - and responsive to the ways in which these interactions will change, as model capabilities improve and human behavior changes. As a starting point, it may be useful to reflect on what’s shared and what’s different across these classes of interactions. Of course, the involved agents themselves - and the process by which they acquire worldly knowledge - can be fundamentally different. Many have pointed out that the scale and content of training data differs across people and AI models. Yet there are high-level structural similarities. Analogous to how a person might listen to conversation between others, LLMs in pre-training learn from traces of interactive reasoning, where the learning agent is not an involved party in the interaction. Even before all the post-training explicitly designed to make LLMs more useful as “interactive” conversational agents, they are exposed to interactions between multiple individuals with distinct epistemic states exchanging information over an unfolding conversation, played out in comment threads and online forums. In this sense, both humans and AI models occupy the role of an observer embedded in a society of interacting agents, uptaking the externalized traces. The form factor of the interaction itself is also worth considering. As noted earlier, all classes of interactions are primarily realized through natural language - a point of commonality. Yet there are also clear differences, e.g. in the availability and latency of responses - a chatbot can answer instantaneously, without incurring a “cost” (in a psycholinguistic, not necessarily monetary, sense) to produce such a response, becoming a sort of assistant on demand. Given this reduction in barriers, one of the most promising, although also potentially perilous, opportunities for interactive LLM use is as a tool for personalized learning. For most of human history, before widespread literacy, learning through the experiences of others was fundamentally interactive: there was no way to learn from another person other than being physically co-located and speaking directly with them. The advent of the written word changed that dynamic; all kinds of knowledge, from mathematics to poisonous mushroom species, became accessible in written, static texts that were broadly accessible. Learning through AI interaction feels like a kind of interpolation between these different modes, in that it’s primarily realized through the familiar form of written text, yet the text is no longer static; it is generated through unfolding interaction with a teacher-like figure. In this sense, learning with LLMs is both an extension of existing paradigms and a qualitatively new phenomenon. Of course, these are very high-level reflections on the structural similarities and differences across different classes of interactions - what would it mean to actually interrogate this rigorously? Forging this path will likely involve revisiting broader and perpetually-unresolved methodological debates on experimental control vs. naturalism, and formal modeling vs. qualitative insight - but one thing that seems clear is the need to design experiments that expose open-ended interaction in the first place, or otherwise aggregate naturally-occurring datasets that reveal traces of interaction. Maybe this seems like kind of a trivial point, but at the same time, much remains to be done in investigating these phenomena through the lens of cognitive science. It feels more attainable—and more important—than ever to study this fundamental aspect of intelligence. Thank you to Bill Thompson for feedback on earlier drafts, as well as to the many people who have helped me develop the ideas here!]]></summary></entry></feed>