MIT: How far can we go down the path of making AI simulate real people through interviews?
Can a large language model become a person's "digital doppelgänger" as long as you chat with them for long enough?
From virtual users and synthetic respondents to personalized AI Agents, an increasingly popular technical path works as follows: first, learn about a real person through interviews, questionnaires and personal profiles, then let the large language model act as this person, and predict how they will answer questions they have never seen before.
But there is a more difficult question here: Even if AI has already learned how you think right now, when placed in a new situation that has never been asked before, can it still answer in a way that is just like you?
To test this, researchers from institutions including MIT proposed HugAgent, which allows large language models to predict the responses of real participants in new scenarios such as universal healthcare, public surveillance and housing densification based on interviews with these participants. The study covered 54 participants and 1,742 evaluation items, and tested models including GPT, Claude, Gemini, DeepSeek, Llama and Qwen. The related paper was selected for EMNLP 2026 Main Oral.
Paper link: https://arxiv.org/abs/2510.15144
Code link: https://github.com/jajamoa/HugAgent
Project homepage: https://jajamoa.github.io/HugAgent/
The study found that personal interviews do help models simulate a specific person more accurately. But a more notable phenomenon is that: The longer you chat with a person, the clearer AI will be about how they think now, but it does not necessarily know better how they will answer when conditions change.
This also brings up a more fundamental question for "digital doppelgängers": Is AI really building a model of you, or is it just creating an increasingly detailed personal profile?
From "like an average person" to "like this specific person". HugAgent further asks: when conditions change, can AI still represent this person?
Why HugAgent is needed
From "being human-like" to "being like this specific person"
In the past few years, simulating humans with large language models has gradually become a formal research direction. Researchers have asked LLMs to act as consumers, voters and experimental participants. Going a step further, they have also started trying to simulate a real existing individual: build a "digital doppelgänger" through long interviews, questionnaires and personal materials, then let AI continue to represent the person themselves in subsequent new questions.
This path sounds very natural. The more information a person tells AI, the better AI understands them, and it seems that the more accurately it can predict their responses.
But there is a critical leap here:
Describing a person is not the same thing as predicting a person.
Suppose a person states in an interview that they support universal healthcare, but at the same time worry about cost, efficiency and personal choice. The model may already be able to accurately judge their current stance.
But what if the conditions change?
If the new plan can save their family 3,000 US dollars a year, or private insurance can still be retained, will they still make the same choice?
At this point, what the model needs to do is no longer just remember "they support universal healthcare", but to judge from past expressions what they really value, and how they will re-weigh when different conditions change.
This is exactly the question HugAgent wants to test:
If we build a sufficiently rich persona through interviews, can it predict how this person will answer outside the interview?
This is an unavoidable question after moving from "being human-like" to "being like this specific person".
HugAgent does not only record answers
It also asks for the "why"
To test whether AI can simulate "a specific person", first of all, we need to know exactly what this person themselves think.
Therefore, HugAgent does not start from public personal profiles or online texts, but recruits real participants again. The team selected three topics with obvious divergences in the United States that also involve real value trade-offs: universal healthcare, public surveillance, and housing densification, which means relaxing some zoning restrictions to allow the construction of taller and denser residences. After quality screening from more than 120 recruits, 54 participants were finally retained.
Step 1: Questionnaire, record "the you right now".
Participants first filled in demographic information, then gave support scores from 1 to 10 on different issues. At the same time, they also needed to evaluate how much a series of reasons influenced their stance.
For example, two people who both oppose housing densification may have totally different concerns: one may be most worried about community style and property value, while the other mainly cares about traffic, rent or housing supply. The final "support or oppose" is just an answer, but what really distinguishes different people is often the reasons behind the answer and how important each of these reasons is.
Together, this information constitutes a person's current belief state, that is, what stances they hold now, and the reasons that support these stances.
Step 2: Change conditions, record "what if...".
Next, the questionnaire began to change one of the conditions.
What if local rents drop by 10% to 15% after housing densification? What if universal healthcare can save a family 3,000 US dollars a year? What if public surveillance footage is only stored for 48 hours?
Participants needed to re-state their stance in each new hypothetical scenario, and re-evaluate the importance of relevant reasons.
In this way, the research team not only knows how a person "thinks now", but also records how they will re-answer when a certain condition changes. The change between the two states before and after constitutes the belief update that HugAgent wants to predict.
More importantly, these answers come from the participants themselves. When the model takes the test later, it will not directly see the correct answers under the target scenario.
Caption: The interview process of TraceYourThinking. The system starts with open-ended questions, dynamically builds a causal belief graph based on the participant's answers, and automatically generates the next round of follow-up questions to gradually understand the relationship between a person's views, reasons and different factors.
Step 3: Think-aloud interview, slowly ask out the "why".
Recording only the questionnaire answers is not enough.
Participants then participated in a semi-structured think-aloud interview through the team's open-source chatbot TraceYourThinking, where they could type or speak directly, stating the reasons for their judgments while thinking.
This is not a fixed list of questions from start to finish.
The robot will dynamically sort out this person's "causal belief graph" based on what the participant said earlier: what factors they mentioned, in what direction these factors will change the stance, and how strong the influence is. Based on this information, the system automatically generates the next round of follow-up questions, and continues to ask about parts that are not clearly stated, contradictory, or worthy of further expansion.
For example, if a person says they are worried about the cost of a certain policy, the robot can continue to ask: If the cost drops, to what extent will this change your attitude? If they say "I'm not sure about that", the system will also follow up on this uncertainty.
"I'm not sure" is information in itself.
It may mean that this person's stance is not firm, or that two values are pulling against each other. Compared with a simple score from 1 to 10, the think-aloud interview leaves richer clues: what this person cares about, why they care, which judgments are certain, and which parts are still hesitant.
Finally, the model gets the background information of this person and this interview, not the real answers given by the person themselves under the target scenario.
Caption: Two types of evaluation tasks of HugAgent. Belief State Inference infers a person's current beliefs based on interviews; Belief Dynamics Update adds new hypothetical scenarios to predict how the same person's stances and reasons will change.
HugAgent then turns these materials into two types of exams.
The first type is Belief State Inference, with a total of 356 questions, testing whether the model can infer the unstated current beliefs of this person from the interview.
The second type is Belief Dynamics Update, with a total of 1386 questions, testing whether the model can predict the person's new stance and reason weights after a new hypothetical scenario is given.
The entire benchmark contains a total of 1742 evaluation items.
To set an empirical upper limit for the task, the team invited some participants to answer the same questions again 14 days later. The consistency of participants' answers with their own answers two weeks ago was 84.84% and 85.66% on the two types of tasks respectively, which was used as the ceiling (empirical upper limit) for subsequent model comparison.
At this point, an originally very abstract problem has finally become directly measurable:
If you chat more and more with a person, will their "digital doppelgänger" become better and better at predicting them?
Core result: It is easy to understand "the you now", but it becomes difficult when conditions change
The most intuitive result of HugAgent is that there is a clear gap between the performance of the model on the two types of tasks.
In the Belief State Inference task, the best performing method achieved an accuracy of 77.56%, and GPT-4o reached 74.66%. When participants re-answered the same question 14 days later, the self-consistency was 84.84%, which the paper used as the empirical ceiling for this task.
For Belief Dynamics Update, the best performing Claude Sonnet 4.5 reached 68.61%, LLaMA 3.3 70B reached 67.57%, and GPT-4o reached 63.11%; the corresponding empirical ceiling was 85.66%.
In other words, under their respective evaluation settings, large language models are already quite good at restoring how a person "thinks now" from interviews, but when facing a new scenario that the person has never answered before, there is a more obvious gap between the model's output and the participant's own answer.
Caption: The overall performance of different models on the two types of HugAgent tasks. Compared with inferring a person's current beliefs, when predicting responses in new scenarios, there is still a more obvious gap between the model and the upper limit of the participant's own consistency.
However, this does not mean that responses in new scenarios cannot be predicted from interviews at all.
Taking GPT-4o as an example, if no personal interview is provided and only demographic and other information is relied on, the accuracy of Belief Dynamics Update is only 39.83%; after adding the full interview of this person, it rises to 63.11%.
Qwen2.5-32B also showed similar changes, rising from 32.12% to 58.96%.
In other words, the interview does not just add a longer context to the model. The words that this person themselves said do contain individual information that predicts how they will answer later.
Further cross-person experiments also got similar results: if the personal interview is replaced with another person's interview, the accuracy of GPT-4o on the update task drops to 39.30%, almost returning to the level without the person's own interview.
Caption: Performance comparison between no personal interview and adding the full personal interview. Personal interviews significantly improve the model's ability to predict specific participants.
Since interviews are useful, a very natural thought is: Then chat more.
If the digital doppelgänger is mainly a matter of information volume, then 5 rounds of interviews should be worse than 10 rounds, and 10 rounds should be worse than the full interview.
But HugAgent did not observe such a stable relationship.
Taking Qwen2.5-32B as an example, as the interview content increases, the accuracy of Belief State Inference rises from 68.84% to 73.02%, and finally reaches 77.17%.
However, Belief Dynamics Update did not improve synchronously, instead changing from 61.79% to 60.73%, and reaching 58.96% under the full interview.
This does not mean that interviews are useless. Previous experiments have shown that the difference between "not knowing this person" and having this person's own interview is obvious.
What is really worth noting is:
It is very important to go from having no personal information to having personal information, but continuously piling up more and more materials about a person will not automatically bring more accurate predictions for new scenarios.