Synthetic Consumer LabSynthetic Consumer LabTry the demo
Article — 5 min read

LLM-Based Doppelgänger Models: Leveraging Synthetic Data for Human-Like Responses

Learning and imitating individual opinions with an LLM-based Doppelgänger model: accuracy, heterogeneity, and single-device feasibility.

Research Questions

  1. Can an LLM-based model learn and imitate an individual’s thoughts and judgments from their conversations?
  2. Can such individual-centred models be used in survey research?
  3. Is it feasible to simulate personalised opinions in a practical single-device environment (e.g., a 40GB GPU)?

Results

  • The Doppelgänger model replicated individual opinions with high accuracy.
  • It achieved 67% accuracy on a 5-point scale and 80% accuracy on a 3-point scale.
  • Longer context windows (e.g., 3,000 tokens) and more training epochs increased accuracy.
  • It outperformed commercial models such as GPT-3.5 Turbo and Gemini.
  • It proved feasible to run on a single 40GB GPU.

Findings

  • The Doppelgänger model can reproduce opinions not only at the group level but also at the individual level with strong fidelity.
  • The paper’s Surveyed LLM (0.46 / 0.65 accuracy) and the commercial models performed poorly at individual-level imitation.
  • Predictions based solely on metadata were weak and biased, underscoring the importance of conversational data.
  • The model learned effectively even with an average of 21 conversation samples per individual.
  • The approach provides a powerful tool for capturing individual-level heterogeneity and generating personalised responses.

Scores

  • LLM Models: 5
  • Synthetic Data: 5
  • Method: 5
  • Speed: 4
  • Ethics: 2
  • Accuracy: 5
  • Demographics: 2

If you would like to explore this research in more detail, click here to read the full paper.

Read the full paper ← All articles