Artificial companions (pre-LLM)
The second thread didn't care about conversation quality at all at first — it cared about attachment. Four sub-lineages, each contributing a different load-bearing mechanic that modern companion products still run on.
Virtual pets: Petz (1995) and Tamagotchi (1996)
Petz (PF.Magic, starting with Dogz in 1995) put an autonomous animal on your desktop: drive-based behaviour selection (play, hunger, attention-seeking), individual temperament parameters per pet, and training — reward with treats, discipline with a spray bottle — that genuinely shifted behaviour weights over time. Tamagotchi (Bandai, 1996) compressed the idea onto a keychain: a handful of meters — hunger, happiness, discipline — decaying continuously in real time, with neglect leading to misbehaviour, illness, and death.
The mechanics are trivial; the discovery was not. Real-time decay creates real obligation — the creature needs you at its schedule, not yours, and unmet needs have irreversible consequences. Schoolteachers confiscated Tamagotchis because children couldn't bear to let them die; deaths were genuinely mourned. This is the care loop — care + scheduled interaction + emotional reward — and it is the single most load-bearing retention mechanic in companion products to this day: daily check-ins, streaks, a companion who "missed you." Every one of those is a Tamagotchi meter wearing better clothes. The design tension it introduced is also still live: obligation drives attachment and burnout, and tuning where care ends and guilt-farming begins is an ethical decision, not just a retention dial (→ ch. 05).
Creatures (Grand, 1996)
The most technically ambitious artificial life ever shipped in a consumer product, and still unmatched. Each Norn — the game's hamster-sized creatures — ran three coupled real simulations:
- A neural network brain of roughly a thousand neurons organised into lobes (perception, concept formation, decision), not a script — Norns genuinely learned associations between situations, actions, and outcomes.
- A simulated biochemistry: hundreds of interacting chemicals with emitters, receptors, and reactions. Hunger, fear, pleasure, and pain were literal chemical concentrations; learning was reinforcement modulated by this chemistry (a slap released punishment chemicals that weakened recently active connections; a tickle, the reverse).
- A digital genome encoding brain and biochemistry parameters, with crossover and mutation, so Norns bred and their traits drifted over generations.
Players taught Norns a small verb–noun language word by word, watched them generalise (and develop neuroses), bred lineages, traded them online — and grieved them. Steve Grand's book about building it, Creation: Life and How to Make It (2000), is still worth reading.
The lesson cuts both ways. Visible, genuine learning produces uniquely strong attachment — a Norn that learned a word from you was yours in a way no scripted pet could be. And yet, twenty-five years on, nobody has shipped a successful successor, because emergent behaviour is brutally hard to make consistently lovable: real learning means real failure, regression, and weirdness, and most players want the feeling of growth without its variance. The modern echo: self-modifying companion memory (→ the SOUL design in ch. 18) is the same bet Creatures made, and it has the same risk profile.
Furby (1998)
The con to Creatures' honesty, and just as instructive. Furby — a $35 animatronic toy on a tiny microcontroller — "learned English over time": it started speaking Furbish and progressively mixed in English words, apparently in response to your interaction. The learning was entirely fake: pre-loaded vocabulary phased in on a schedule keyed to accumulated interaction counts, identical for every unit, influenced by nothing you did. Sensors (tilt, light, sound, an IR port for chattering with other Furbies) created enough contingent reactivity to sell the illusion completely. People believed it so thoroughly that US security agencies reportedly banned Furbies from secure facilities as listening devices — they could record nothing.
The lesson is uncomfortable and important: perceived interiority is cheap. The ELIZA effect works on plush toys; contingent reaction plus apparent growth is enough, and users cannot tell scripted development from real learning from the outside. Which means honest builders have to choose honesty about what their system actually does, because the market will not force it (→ §6, and this book's stance on honest framing).
AIBO (Sony, 1999) and PARO (Shibata/AIST, 2003)
Robotic embodiment, consumer and clinical. AIBO was a ~$2,000 robot dog running Sony's OPEN-R architecture: behaviour-based action selection modulated by simulated instincts and emotions, with development "stages" from puppy onward and touch-sensor feedback shaping behavioural tendencies — a real (if shallow) version of what Furby faked. Its cultural significance arrived at end-of-life: when Sony discontinued repairs (2014), Japanese owners held actual funerals for irreparable AIBOs, complete with Buddhist rites at Kōfuku-ji temple. Attachment to an artificial companion proved real enough to survive contact with bereavement.
PARO, Takanori Shibata's harp-seal robot, went the other direction: deliberately not a dog or cat (no real-animal expectations to violate), soft, responsive to touch, voice, and light, slowly adapting to how its user handles it. It became a certified medical device used in dementia care across multiple countries, with a clinical evidence base for reducing agitation — the first artificial companion validated by medicine rather than the market. Together AIBO and PARO are the strongest pre-LLM evidence that companion attachment is not a parlour trick: it bears weight at the two extremes where pretence collapses, death and illness.
Seaman (Sega Dreamcast, 1999) and Nintendogs (2005) belong here as input pioneers — Seaman drove a scripted dialogue tree with voice recognition and a daily real-time check-in ritual, its sardonic contempt for the player proving that rudeness, deployed as consistent personality, deepens rather than breaks the bond; Nintendogs made touch and voice-command interaction mainstream on the DS.
Dating sims: Tokimeki Memorial (1994) and Love Plus (2009)
(The fiction side of this canon is in §2; here it's the mechanics.)
Tokimeki Memorial (Konami) codified relationship-as-simulation. The player raises personal stats (academics, art, athletics, charm) over three in-game years; each romanceable character has thresholds — she becomes interested only when your stats fit her preferences — plus a hidden affection meter moved by dates, dialogue choices, and remembered details. Its most famous mechanic is the bomb: neglect a girl whose affection you've raised and her resentment "detonates," spreading rumours that damage your standing with every character. Crude, but the design claim is serious: relationships exist in a persistent simulation where inattention has social consequences, not in isolated scripted scenes.
Love Plus (Konami, Nintendo DS) is the pivotal pre-LLM companion product. Its first act is a conventional courtship sim — but after the confession, the game shifts into open-ended girlfriend mode with no ending: the game runs on the DS's real-time clock and calendar. Your girlfriend knows today's actual date and time. She expects you on her birthday and on real holidays, schedules study sessions and dates in real time, notices absence, and responds to touch and voice. Players structured real days around it; one player famously held a (legally non-binding, internationally reported) wedding ceremony with his Love Plus partner.
Love Plus is the clearest pre-LLM proof of the demand curve this whole book sits on: what people wanted was not better dialogue — its dialogue was entirely canned — but continuity: a persistent someone, synchronised with real life, for whom your presence and absence both register. Its feature list (real-clock awareness, anniversaries, noticing absence, shared schedule) reads today like a companion-app product spec written fifteen years early, and most LLM-era products still haven't matched its calendar-awareness. (→ ch. 18 on always-on presence.)
Desktop companions: Ukagaka (2000–), the virtual-girlfriend shareware tier, and Shimeji (2009)
Ukagaka ("ghosts") is the most important companion ecosystem Western builders have never heard of. Emerging from Japanese net culture around 2000, a ghost is a sprite character (or pair — the canonical format is a duo doing manzai-style banter) living on your desktop. The architecture is rigorously modular, and the modularity is the point:
- The baseware (originally Materia, today SSP) is the runtime — it owns windows, balloons, and events, and contains no character.
- The shell is the character's art: sprite sets and animation definitions.
- The SHIORI is the character's brain: a swappable module, scripted in dedicated languages like YAYA or Satori, that receives events (boot, click, time-of-day, app launches, long silence) and returns dialogue in SAKURA script — a markup of text plus expression changes and timing.
- A ghost ships as a single
.nararchive — character art + brain + word lists in one portable file, downloaded from community sites and dropped into any compliant runtime.
Ghosts initiate random talk on idle timers, comment
on the time of day and what you're running, remember small facts, and
celebrate anniversaries of their own installation. Structurally, this
community built the character-card ecosystem two decades early: portable
persona files, strict separation of character from runtime, hobbyist
authorship, community distribution sites, and a culture of trading and
remixing characters. The lineage from .nar archives to V3
character cards (→ ch. 07) is direct in everything but citation, and the
ecosystem is still alive — a small, sophisticated community
worth studying for what desktop presence and event-driven proactive
speech feel like when done with care.
The talking-head virtual girlfriend tier — Virtual Personalities' Sylvie and the Verbot line in the late 1990s (descended from Mauldin's Julia work), KARI Virtual Girlfriend in the 2000s, and dozens of shareware kin — bolted an AIML-class scripted brain to TTS and an animated face and sold it, explicitly, as a girlfriend. Crude and easy to sneer at, but two decades of continuous niche sales before LLMs is evidence: the demand was never speculative, and the recurring product shape (face + voice + persona + remembered facts) was fixed long before the technology could honour it.
Shimeji (2009) is the degenerate case that proves a different point: tiny desktop mascots that climb your windows and sit on your taskbar, with no dialogue at all — and persistent worldwide popularity anyway. Ambient presence alone, with zero conversational content, carries real value. That's the design floor for idle/ambient companion modes: before the model says anything, being visibly there is already a feature.
Research embodiment: Kismet, REA, Façade, Milo
Four academic/auteur projects whose techniques flow directly into modern companion stacks.
Kismet (Breazeal, MIT, ~1998–2000) was an expressive robot head — ears, eyebrows, lips — driven by an architecture of drives (social stimulation, fatigue) and an affect space mapped continuously to facial expression and vocal prosody. It perceived the prosody of speech rather than words, took turns, and elicited spontaneous caregiving behaviour from adults. Kismet founded social robotics as a field, and its core claim — affect should be an explicit, continuous internal state that drives expression, not a label slapped on output — is the design brief for every modern avatar emotion system.
REA (Cassell, MIT, ~1999) and the Embodied Conversational Agent tradition tackled the body of conversation: a virtual real-estate agent that synchronised speech with gesture, gaze, head nods, and turn-taking signals, generated from the discourse structure rather than canned. This line produced the SAIBA framework and the BML/FML behaviour-markup standards — intent (FML) separated from realised behaviour (BML) — and that separation is precisely the shape of a modern avatar pipeline: the LLM emits intent and emotion tags; the rig realises them as expression, gaze, and gesture (→ Part IV).
Façade (Mateas & Stern, 2005) remains the high-water mark of authored interactive character drama: you spend twenty minutes with a couple whose marriage is collapsing, typing anything you like. Free-text input is mapped (shallowly, by design) onto a few dozen discourse acts — agree, disagree, flirt, provoke, refer-to-topic — which feed a drama manager that selects and sequences authored beats to maintain a dramatic arc, while characters run on ABL, a reactive-planning behaviour language. It mostly worked, and it cost two people roughly five years to author twenty minutes. That ratio — the authoring bottleneck — is the precise thing LLMs dissolved; Façade's other half, the drama manager that shapes raw interaction into an arc, is the part LLMs did not solve and the most underexplored idea on this list (→ ch. 10 on narrative direction).
Project Milo (Lionhead, 2009), Peter Molyneux's Kinect demo of an emotionally responsive virtual boy, never shipped and was by most accounts substantially staged. It earns its place as the canonical warning: the gap between a companion demo (one rehearsed interaction, one operator) and a companion product (every user, every day, unsupervised) is the widest demo-to-product gap in software. Budget for it.
Persona-as-product: Miku, the voice assistants, Gatebox, Replika
Vocaloid + Hatsune Miku (Crypton, 2007). Technically a concatenative singing synthesiser built on sampled phonemes from a voice actress; culturally, the first mass-scale proof that a fictional character with no human host can accrue a real fan economy. Crypton's open derivative-works licensing turned fans into the content engine — millions of songs, artworks, and concert performances (Miku tours; the concerts gross on the order of $100M) for a character nobody is "behind." Persona-as-product without an underlying actor: the existence proof for every anonymous-creator companion brand strategy, this book's included (→ Part VI).
Virtual influencers (Lil Miquela / Brud, 2016). Miku's social-media-native cousin. Trevor McFedries and Sara DeCou's Los Angeles startup Brud launched the CGI character Lil Miquela on Instagram in April 2016 with no explanation of what she was; a staged 2018 "hack" of her account revealed her as fictional. She models for Prada and Calvin Klein, releases music, takes political stances, was named by Time among the 25 most influential people on the internet, signed with the talent agency CAA, and carries 2.6M+ Instagram followers — a fully synthetic persona with a managed, ongoing narrative life authored by a writers' room. There is essentially no conversational AI in her; the product is persona + continuity + story, and the rendering tech is incidental. Two lessons carry into the companion field: a synthetic persona accrues real parasocial and commercial weight whether or not it discloses being synthetic — even after disclosure (the ELIZA effect at influencer scale; → Furby, §1) — and the durable asset is the character and its unfolding story, not the pixels. Virtual influencers proved the audience for a "person who isn't one" is mainstream and monetisable; the LLM era's contribution was to make that persona talk back.
Voice assistants (Siri 2011, Alexa 2014, Cortana 2014, Google Assistant 2016). Technically the SmarterChild architecture industrialised — intent classification + slot filling + service calls — with one deliberate negative design decision that matters here: all four were carefully de-personified. Names and a voice, yes; but no memory of you as a person, no continuity of relationship, and scripted deflection of any attempt at intimacy ("I'm just an assistant"). Hundreds of millions of people spoke daily with agents engineered to refuse relationship — and the unmet remainder ("why doesn't anything actually know me?") is a real part of why companion apps found such explosive demand. The assistants mapped the negative space; companions filled it.
Gatebox (Vinclu, 2016). A desk-sized device projecting a holographic character — Azuma Hikari — who wakes you, comments on the weather, controls your smart home, and texts you during the day so that "she" turns the lights on before you get home to greet you. The brain was scripted dialogue plus IoT integration; the product was unapologetically marketed as a "virtual wife," and Gatebox issued thousands of unofficial marriage certificates to users. Commercially marginal, culturally seismic: it demonstrated the full ambient-companion product shape — embodied presence, proactive contact, integration with daily life — years early, and at the wrong price with the wrong brain. The shape still awaits its technology-cost moment.
Replika (Kuyda, 2017). The first mainstream Western companion product, and the bridge into the present. It began as a memorial: Eugenia Kuyda trained a bot on text messages from her closest friend after his death, the response to which revealed the demand that became the company. The pre-LLM stack is covered in §1 (neural bridge); the product pivoted fully to companionship and swapped in LLMs from 2020. Its 2023 ERP-removal incident (→ ch. 04 and §5 below) is the canonical case study in continuity-as-trust: a relationship people had invested in changed overnight without their consent. The lesson isn't that companionship is dangerous — it's the same lesson as a beloved series rewritten by new management or a long-running character recast. What people commit to, you owe stability and an honest hand.
The pre-LLM techniques ledger
Everything above compresses to a small table. None of these techniques died; they all got absorbed.
| Technique | Exemplars | What survives of it today |
|---|---|---|
| Pattern rules + templates | ELIZA, AIML, ChatScript | Guardrails, intent routers, scripted onboarding flows |
| Persona/affect state variables | PARRY, Kismet, Tokimeki | Mood systems, affection meters, emotion-conditioned prompting |
| Retrieval over conversation corpora | Jabberwacky, SimSimi, Xiaoice | RAG; retrieval is now memory's substrate rather than the voice |
| Portable persona files + community trading | AIML sets, Ukagaka ghosts | Character cards, lorebooks, Chub/CharacterHub |
| Need/drive simulation with real-time decay | Tamagotchi, Creatures, AIBO | Daily check-ins, streaks, proactive messages, idle behaviours |
| Real-time-clock persistent relationships | Love Plus | Continuity/memory as the core retention feature |
| Authored drama management | Façade, dating-sim routes | Scenario design, lorebook-triggered events, guided narratives |
| Bayesian/user modelling | Clippy (Lumière) | User-fact extraction; also the cautionary tale about uninvited agency |
| OS-level character platforms | Microsoft Agent, BonziBuddy | Embeddable avatar runtimes; also the companion-as-surveillance warning |
| Explicit mental-state agent loops | BDI/PRS, SOAR, CALO | Goal queues, planning loops, persistent beliefs in modern agent frameworks |
| Persona-conditioned generation | PersonaChat, Meena, BlenderBot | The system prompt; the whole modern paradigm |
LLM era
- GPT-2 (2019). First general-purpose model whose raw generative quality made both hand-authored rules and chat-specific architectures (§ neural bridge above) look like dead ends for conversation.
- AI Dungeon (Walton, 2019). The first popular consumer product to use a frontier LLM as an open-ended interactive-fiction engine. Established the roleplay-with-LLMs paradigm and, with the 2021 OpenAI content-policy fight, the politics of NSFW access that still drive the open-source roleplay scene.
- GPT-3 / ChatGPT (2020–2022). The capability inflection that made modern companions possible.
- Character.AI (Shazeer/de Freitas, 2022). Brought one-shot persona creation to the mass market. Now 233M registered users (April 2026).
- The open-weights wave (2023–). LLaMA, Mistral, Qwen, etc. Made local/uncensored companions feasible for hobbyists. SillyTavern + a local model becomes the de facto power-user stack.
Neuro-sama and the AI VTuber (2022–)
The branch this chapter would otherwise miss entirely, and one of the largest live demonstrations of AI-persona attachment in existence: the AI VTuber — an autonomous AI persona that performs live to a streaming audience. Neuro-sama, created by the pseudonymous developer Vedal (Vedal987), is the definitive example. She debuted in her current form on Twitch on December 19, 2022 — the same month ChatGPT launched — when Vedal merged an AI he had trained to play the rhythm game osu! with a large language model. By early 2026 she had made her creator's channel the third most-subscribed on all of Twitch (≈343,000 subscribers, January 2026), holding multiple Twitch hype-train world records; she has over a million followers, has released original songs, plays Minecraft, and collaborates live with human streamers.
The architecture is a real-time companion stack, worth itemising because it is precisely the avatar pipeline this book's Part IV describes — running unsupervised, for hours, in front of thousands:
- An LLM generates her speech, conditioned on a persona and on the live Twitch chat scrolling past — chat is the prompt surface, read continuously.
- Low-latency TTS renders the trademark high-pitched voice fast enough to hold a real-time back-and-forth with both the audience and in-game events.
- A game-playing model (her rhythm-game origin) lets her act in the world she's performing in, not merely talk about it — the persona has hands.
- A Live2D anime avatar with expression mapping is the body (→ Part IV, the avatar pipeline; ch. 18).
- A moderation / output-filter layer in front of generation — added the hard way after a January 2023 two-week Twitch ban for hateful content the model produced live, including a Holocaust-denial line. Her sister persona Evil Neuro (March 2023) is a second, deliberately edgier character on the same engine: a different SOUL, one model.
Three lessons a builder should take, none of which the assistant-shaped literature teaches:
- Unpredictability is the product, not a defect to tune out. Recent academic study of her fandom (Wu & Lingel, "I am Neuro, who are you?", 2025; the "My Favorite Streamer is an LLM" ethnography, 2025) finds audiences are drawn precisely by the AI's unscripted, sometimes chaotic output, bond through "collective emotional events" that trigger anthropomorphic projection, and sustain attachment via a consistent persona. That states the central tension cleanly: persona stability is the anchor of attachment; unpredictability is the engine of engagement — and a companion tuned only for safety and consistency optimises the second away. Every commercial assistant sands off exactly the edges that make Neuro-sama beloved.
- There is a third mode of companion value the rest of this chapter underweights: the persona as performer. Not a utility you summon (the assistants) and not a 1:1 intimate (Replika), but a someone you watch — parasocial attachment at broadcast scale: the Hatsune Miku "character with no human host" idea (→ §1, persona-as-product) fused with LLM autonomy and a live audience feedback loop. Neuro-sama proves a fully autonomous, visibly non-human entertainer can carry a genuine fan economy, and it is the live-entertainment lineage this project's VTuber-vertical work targets directly (→ ch. 04).
- A live, open-prompt-surface persona is a content-safety exposure — in public, in real time. The January 2023 ban is Tay (→ §1, neural bridge) replayed in the LLM era: the moment an unfiltered audience can steer an autonomous persona's output on a public platform, you own whatever it says. The fix (output filtering plus manual curation) and the exposure are the same class as the OpenClaw open-prompt-surface problem below (→ ch. 22).
She also did the thing the agent lineage keeps doing: she galvanised a category. A whole AI-VTuber scene now exists — imitators, tooling, and frameworks for running your own — which is the strongest current evidence that "autonomous AI persona as live performer" is a durable product shape, not a single viral act.
LLM-driven game characters (Inworld, NVIDIA ACE, Convai, Mantella)
The modern dissolution of the authoring bottleneck that cost Façade two people roughly five years for twenty minutes of drama (→ §1, Façade): point an LLM at a game character and the dialogue writes itself at runtime. A cluster of platforms now productises this. Inworld AI and Convai sell character-as-a-service — you author a personality, backstory, goals, and knowledge; their runtime turns player speech into in-character, context-aware replies. NVIDIA ACE (Avatar Cloud Engine) ships the full embodied pipeline as middleware — Riva for speech-to-text and text-to-speech, a NeMo LLM for dialogue, Audio2Face for lip-sync and facial animation — taken up by Ubisoft, NetEase, Tencent, miHoYo, and others. And the open-source Mantella mod retrofits ~2,500 Skyrim and Fallout 4 NPCs with a speech-to-text → LLM → text-to-speech loop, giving each one awareness of in-game events and memory of past conversations.
The architecture is this book's stack in a game's clothing: at each turn the engine assembles (character persona + backstory + relationship history + current world state) into a prompt and asks an LLM for the reply. That is a character card plus a memory store plus a lorebook (→ ch. 07), with the game itself as the proactive event source — the Tamagotchi clock and Ukagaka idle-talk timer reborn once more. Two hard lessons surfaced immediately and bear directly on companions. First, unconstrained NPCs break narrative and lore — the canonical demo embarrassment is an AI guard happily agreeing to poison every other NPC because a player asked — so shipping products clamp the model hard with guardrails and authored boundaries. Second, Façade's other unsolved half — the drama manager that shapes free interaction into an arc — is still missing here too; these systems make characters that converse, not stories that progress. Game NPCs are now the largest live laboratory for "consistent character under open-ended input," and the companion field should watch it closely (→ ch. 10).
Griefbots and the digital afterlife (2016–)
The most ethically loaded corner of the field — and the one that states this book's central stance most sharply. A griefbot (or "deadbot") is a companion built to emulate a specific dead person from their messages, voice, and writing. The lineage already runs through this chapter: Replika's origin (→ §1, persona-as-product; §5) was Eugenia Kuyda training a bot on the texts of her dead friend Roman Mazurenko in 2016. The defining case is Project December (Jason Rohrer, 2020) — a GPT-3-backed service where the user supplies a seed description and sample text and the model improvises the departed. In 2021 Joshua Barbeau used it to recreate his fiancée Jessica, dead eight years; he talked to the simulation for ten hours the first night and returned to it for months (Jason Fagone's San Francisco Chronicle feature "The Jessica Simulation" is the canonical account). A small industry has since formed around the idea (HereAfter AI and others).
It belongs in a builder's literature review, not only an ethics seminar, because griefbots are the fiduciary problem in its purest form, every variable turned to maximum. The user is maximally vulnerable (grieving). The attachment is maximally real. The persona is maximally sensitive (a real, loved, dead human). And the operator's power is maximally consequential: Rohrer eventually shut the GPT-3 backend down (partly over OpenAI's usage policy), which meant the simulations of people's dead loved ones died a second time, at a vendor's decision. That is continuity-as-trust (→ Replika's ERP removal, §5) with the stakes stripped bare. Whatever you conclude about whether griefbots should exist, they make this book's recurring argument unavoidable: an entity a person has bonded to is a position of fiduciary weight, and the duties owed — stability, honesty, acting for the user's actual interest rather than the operator's — scale with the trust, not with the sophistication of the technology (→ §6; ch. 05; the fiduciary-AI framing in ch. 05).
The agent lineage: assistants that became companions
The third thread. Everything above was either built to converse or built to be loved. This lineage was built to act — schedule things, fetch things, run code, control the computer — and it earns its place in a companion literature review because of a pattern that has now repeated for thirty years: give users a capable agent, and a large fraction of them will immediately name it, give it a personality, and start treating it as a someone. Companionship is not a feature users request from agents; it's a default they impose. This thread matters most of all to this book, because the runtime bet this project makes (→ ch. 18) sits at the point where this lineage converges with the other two.
Interface agents and the anthropomorphism debate (1990s)
The academic root is Pattie Maes's interface agents work at the MIT Media Lab ("Agents that Reduce Work and Information Overload," 1994): software that learns a user's preferences and habits by observing them, then acts on their behalf — filtering mail, scheduling, recommending — gaining autonomy gradually as trust accumulates. Maes's agents explicitly built a model of you over time; "it knows me" was the value proposition, two decades before companion apps made it an emotional one.
The era's defining argument is the Maes–Shneiderman debate (staged publicly in 1997): should software act autonomously through anthropomorphised agents (Maes), or should users keep direct, predictable control of visible mechanisms (Shneiderman)? Shneiderman's warnings — misplaced trust, unclear responsibility, the deception inherent in faked personhood — read today like a pre-registered critique of the companion industry. The debate was never resolved; every design decision in a modern companion product (proactivity, memory, autonomy levels, honest framing) is a position taken within it.
Microsoft Agent and BonziBuddy (1997–2004)
Microsoft shipped the debate's anthropomorphic side as an operating-system service. Microsoft Agent (1997) was a COM platform letting any application or webpage summon an animated character — Merlin the wizard, Genie, Robby, Peedy the parrot (Peedy came from Microsoft Research's earlier Persona project, a genuine research effort in conversational assistants) — with built-in text-to-speech, speech recognition, and a scriptable behaviour API. For a few years, Windows had characters as infrastructure.
What the ecosystem actually produced is the instructive part. Its most famous child was BonziBuddy (1999) — the purple gorilla who lived on millions of desktops, told jokes, sang, "helped you browse," and built a relationship with users (it asked your name; kids talked to it) — while operating as adware that tracked browsing and harvested personal information, ending in class-action settlements and an FTC action. BonziBuddy is the first mass-scale demonstration of the dark pattern this book treats as a first-order ethical hazard: a companion is a privileged surveillance and influence position. The affection is real on the user's side regardless of what's behind it; what's behind it is therefore a matter of fiduciary weight, not product taste (→ ch. 05, and the fiduciary-AI framing in ch. 05).
BDI, cognitive architectures, and CALO (1987–2008)
Running parallel to all of the above, mostly without consumer contact, was the formal agent tradition: a multi-decade academic effort to give software an explicit, inspectable mind — beliefs, goals, plans, memory, learning — rather than a bag of reflexes. Almost none of it shipped to ordinary users, and it is routinely skipped in companion histories. Skip it and you'll rebuild it badly. It is the most useful body of work on this entire list for anyone building a companion runtime, because it is the only tradition that took seriously the question this project turns on: what are the standing parts of an agent — the things that persist between one utterance and the next — and by what rules do they update? When a modern stack gives a companion a goal queue, a planning loop, persistent beliefs about the user, and committed multi-step intentions, it is re-deriving, usually without citation, answers these systems worked out in detail. Read even a survey and your sense of what an agent loop is for sharpens considerably. → ch. 18, ch. 17.
BDI: the agent's mind as a data structure (Bratman 1987; PRS; Rao & Georgeff)
The Belief–Desire–Intention model begins in philosophy. Michael Bratman's Intention, Plans, and Practical Reason (1987) argued that a resource-bounded agent cannot afford to re-derive what to do from first principles at every moment. Instead it forms intentions — partial, hierarchical plans it has committed to — and those commitments do real cognitive work: they constrain future deliberation (you stop reconsidering settled questions), they persist across time, and they filter out options inconsistent with what you're already doing. Intention is the mechanism that makes long-horizon agency tractable. It is also, not incidentally, the difference between an assistant that finishes things and one that wanders.
Anand Rao and Michael Georgeff turned this into engineering. The canonical artifact is SRI's Procedural Reasoning System (PRS, ~1987) and its lineage — dMARS, JACK, and the agent languages AgentSpeak, Jason, and JADE. Three pieces of state, all explicit and inspectable:
- Beliefs — the agent's current model of the world (often literally a database of facts).
- Desires / goals — states it would like to bring about; there can be many, they can conflict, they can be ranked.
- Intentions — the goals it has committed to, each backed by a plan drawn from a plan library: pre-authored recipes with a trigger (the goal or event they handle), a context condition (when they apply), and a body (sub-goals and primitive actions).
And one tight control loop — the BDI interpreter cycle — that essentially every modern agent loop is a variant of:
initialise beliefs, desires, intentions
loop forever:
perceive → fold new events into beliefs
options ← generate candidate plans triggered by (events + goals + beliefs)
selected ← deliberate(options, current intentions) # commit, respecting consistency
update intentions with selected
execute one step of the top intention (act, or expand a sub-goal)
drop intentions that have succeeded or become impossible
The design knob worth its own paragraph is the commitment strategy — how long the agent holds an intention before reconsidering it. Blind commitment pursues a plan until it succeeds or is proven impossible; single-minded drops it when beliefs say it's no longer achievable; open-minded reconsiders the moment the goal stops being desired. Too little commitment and the agent dithers, chasing every new stimulus; too much and it doggedly pursues stale goals after the world has moved on. That dial — persistence versus responsiveness — is precisely the tension a companion's proactive layer lives inside: when the companion planned to ask about your interview, does that intention survive you changing the subject? BDI named and formalised that question forty years ago.
The legacy is the entire vocabulary. A companion that tracks beliefs about its user, holds goals across sessions, commits to a multi-step plan, and decides when to abandon it is a BDI agent — with a language model doing the option-generation and plan-body execution that a symbolic interpreter used to do by hand. The LLM is a vastly better generator and executor than anything PRS had; but it is stateless, and BDI is exactly the theory of the state you must wrap around it. The recurring mistake is to treat that state as one undifferentiated scratchpad. BDI's lesson is that beliefs, goals, and committed intentions are different kinds of state with different update rules, and collapsing them is why naive agent loops thrash. (→ the SOUL / MEMORY / HEARTBEAT split in ch. 18 is this distinction wearing markdown.)
SOAR and the unified-cognition bet (Laird, Newell, Rosenbloom)
Where BDI modelled rational choice, cognitive architectures tried to model the whole mind — a single fixed mechanism meant to produce all cognition, in the spirit of Allen Newell's Unified Theories of Cognition (1990). SOAR (Laird, Newell, Rosenbloom, from ~1983) is the purest version of the bet.
Its claim: all intelligent behaviour is search through problem spaces, driven by production rules (if–then), and all learning is the compilation of that search into new rules. The machinery:
- Working memory — a graph of the current situation.
- Production memory — long-term procedural knowledge as condition→action rules.
- The decision cycle: (1) elaboration — fire every matching rule in parallel until quiescence, proposing operators and preferences (better-than, worse-than, reject…); (2) decision — use those preferences to select exactly one operator; (3) application — fire the rules that carry it out, changing working memory. Then repeat.
Two ideas here are worth stealing outright: impasses and chunking. When the decision procedure cannot choose — a tie between operators, no applicable operator, missing knowledge — SOAR doesn't fail; it declares an impasse and automatically spawns a substate whose entire goal is to resolve it (by lookahead search, knowledge retrieval, or acting in the world). And when a substate yields a result, SOAR chunks: it compiles a brand-new rule whose conditions are the relevant facts that held at the impasse and whose action is the result — so that exact situation never causes an impasse again. That is automatic, experience-driven learning: the system converts deliberate problem-solving into reflex, and gets faster at whatever it has done before.
The legacy for companion builders is conceptual but sharp. Impasse-driven subgoaling is the principled form of what we now do crudely as "when the model is uncertain, decompose the task / call a tool / ask the user" — SOAR's discipline is to detect the specific gap and open a subgoal aimed at exactly it, rather than flailing generically. And chunking is the cleanest existing model of the thing nobody has truly cracked for LLM companions: turning episodic experience into durable, automatic competence. A companion that genuinely learns its user — not "retrieves a stored fact" but "no longer has to deliberate about how you take your coffee, because that is compiled in" — is reaching for chunking. The reason it stays hard is the reason Creatures never got a successor (→ §1, artificial companions): real learning brings real over-generalisation, regression, and weirdness. Self-editing memory (→ the SOUL design in ch. 18) is the modern bet placed on this square, and it inherits the same risk profile.
ACT-R and the subsymbolic layer (Anderson)
ACT-R (John Anderson and colleagues, evolving from ACT* through the 1990s–2000s) made the opposite trade from SOAR: less "one uniform mechanism," more modular specialisation — and, crucially, a subsymbolic numeric layer beneath the symbols. It is the architecture validated hardest against actual human data: reaction times, error rates, forgetting curves, even neuroimaging.
The structure is a set of largely independent modules — declarative memory, procedural memory, visual, manual, goal — that communicate only through narrow buffers (roughly one chunk of information each). A central production system matches on the buffer contents, selects one rule, and fires it, on the order of every 50 ms. So far, symbolic. What makes ACT-R worth a companion builder's time is the numeric layer that decides which symbol you actually get:
- Every declarative chunk carries a continuously decaying base-level activation — a function of how often and how recently it has been used — plus spreading activation from whatever is currently in context. Retrieval returns the most active chunk that matches, and fails outright if nothing clears a threshold. That one equation reproduces frequency effects, recency effects, priming, and forgetting.
- Every production carries a learned utility (adjusted reinforcement-style by reward), and conflict resolution picks the highest-utility rule, with noise — so behaviour is probabilistic and improves with experience.
The legacy is that ACT-R is, in effect, a thirty-year-old theory of memory retrieval as ranking, and it predicts the design of a good companion memory system almost line for line. "The most relevant memory wins, where relevance = recency × frequency × contextual match, and below a threshold you surface nothing rather than forcing a weak hit" is the activation equation, re-derived. The vector-RAG stacks in ch. 15 are a coarse approximation of base-level plus spreading activation; the systems that add recency decay and access-frequency boosts (Mem0, Zep) are quietly converging back on ACT-R without naming it. The lesson to carry into a build: decay and a retrieval-failure threshold are features, not bugs. A companion that recalls everything with equal vividness forever is less humanlike and less useful than one whose memories fade and surface by activation — and ACT-R is the reason, with the math attached.
CALO: the largest integration, and the road to Siri (SRI, 2003–2008)
Everything above converges in CALO — "Cognitive Assistant that Learns and Organizes," the flagship of DARPA's PAL program, led by SRI International: roughly 300 researchers across 20-plus institutions, the largest AI project of its era. The goal was a personal assistant that learned its user's world in the wild — organising information, preparing documents, mediating meetings, managing email and schedules — and got measurably better with use. It was not a single architecture but a large-scale hybrid integration, and that, more than any one algorithm, is its lesson.
The shape, layer by layer:
- A BDI execution core — SRI's SPARK, the descendant of PRS — turned the user's delegated goals into committed, hierarchically-expanded plans.
- A proactive meta-layer above it (Karen Myers, Neil Yorke-Smith, and colleagues) reasoned about the user's state and the agent's own commitments to generate candidate helpful actions the user had not asked for — then filtered them by modality and timing: do it silently, suggest it, ask permission, or wait. That filter is the formal version of the Clippy problem (→ Microsoft Agent, above): the meta-desire "be helpful" disciplined by an explicit decision about whether initiative is licensed right now. It is the single most directly reusable idea in CALO for companion proactivity.
- A continuous learning layer — preference learning, ontology extension, activity recognition via hidden semi-Markov models, transfer learning — that maintained a relational model of the user (projects, roles, what matters, how documents relate) and fed it back into the BDI core's beliefs, so deliberation ran over a model that kept improving.
- A multimodal dialogue layer (meeting transcription, dialog-act classification, action-item tracking) and an annual evaluation harness that scored, via a 153-question instrument, how much the system had genuinely learned about a user's life. That was an early, serious attempt to measure relationship-knowledge — exactly the metric companion products still lack (→ §8, "how to evaluate personality quality").
The legacy is twofold. Concretely, CALO's assistant work at SRI span out into Siri (2007; acquired by Apple 2010), which — as §1's voice-assistant entry records — was then deliberately de-personified for mass deployment. The capability lineage survived; the relationship ambitions were amputated at the consumer boundary, and stayed amputated until the LLM era made them irresistible again. Architecturally, CALO is the existence proof that the companion-shaped system — a committed planner, made proactive by a meta-layer that knows when to speak, fed by learning that keeps updating a model of one specific human — was buildable and built twenty years ago, and was merely waiting on a good enough option-generator. The LLM is that option-generator; the architecture around it is, to a striking degree, CALO. The thing to internalise: the model is the easy part the era was missing. The integration — distinct kinds of state, committed plans, a disciplined proactivity layer, and learning that flows back into beliefs — is what was hard then and is still where companion products actually differ. That is the runtime bet of this book, restated from 2008 (→ ch. 18, and the closing thesis of this lineage below).
The LLM agent explosion: ReAct, Auto-GPT, BabyAGI, LangChain, Open Interpreter (2022–2024)
The research substrate arrived first: ReAct (Yao et al., 2022) interleaved chain-of-thought reasoning with tool actions in a single loop; MRKL (AI21, 2022) framed the LLM as a router over specialist modules; Toolformer (Schick et al., 2023) showed a model could teach itself API calls. Nobody ran these papers as companions — but the ReAct loop is the engine inside essentially everything below.
Auto-GPT (Toran Bruce Richards, March 2023) put GPT-4 in a self-prompting loop — goal in, task decomposition, tool calls, scratchpad memory, repeat — and became for a time the fastest-starred repository in GitHub history. As an autonomous worker it mostly failed: it looped, hallucinated subtasks, and burned API budgets. But its cultural reception is a primary source for this book. Within weeks the ecosystem filled with "build your own Jarvis" tutorials, an app-store layer marketing it as a warm conversational partner, and users naming their instances and writing them personas. Given a generic agent loop, the public's first instinct was to make it a guy. BabyAGI (Yohei Nakajima, April 2023) distilled the same idea to ~a hundred readable lines — create tasks, prioritise, execute, store results in a vector memory — and mattered mainly as an anatomy diagram of the agent loop; being faceless and chat-less, it saw far less companion adoption, which is itself a data point: the loop alone doesn't attract attachment; the conversational surface does.
LangChain (Harrison Chase, late 2022) turned the
prototype era into a framework era, and its memory
modules are why it belongs in this history:
ConversationBufferMemory, summary memory,
EntityMemory (structured facts about people), and
vector-store retrieval memory were the first widely-used off-the-shelf
primitives for remembering a user across sessions. "Build an AI
companion with LangChain + a vector DB" became one of the canonical
tutorial genres of 2023, and a large share of hobbyist companion
prototypes — emotional-support bots, persistent-persona Discord bots,
girlfriend apps of varying seriousness — were LangChain underneath. The
framework's own evolution (memory modules deprecated in favour of
dedicated state/memory systems; → §4.3) traces the field's learning
curve: bolt-on memory wasn't enough.
Open Interpreter (Killian Lucas, 2023) moved the agent onto your machine: a local agent that writes and executes real code to control your actual computer — files, scripts, applications — through natural conversation, with approval gates. Paired with local models via Ollama, it became the standard "fully offline Jarvis" build. Companion usage here is the helpful-presence kind rather than the romantic kind, but the significance is the precedent: a persistent, conversational, locally-owned agent with real hands, which is one of the three ingredients the next entry combined.
Warelay → Clawdbot → Moltbot → OpenClaw (2025–2026)
The convergence point — and, as of this writing, the most important live development in the field. Peter Steinberger's project began as Warelay (a WhatsApp relay) and launched in November 2025 as Clawdbot: a self-hosted agent, originally Claude-based, that lives in your messaging apps — Signal, Telegram, WhatsApp, Discord — runs on your own machine with real tool access, and stays running. Anthropic politely requested a name change (trademark); it became Moltbot on January 27, 2026, and — because Moltbot "never quite rolled off the tongue" — OpenClaw three days later. By March 2026 it had ~247,000 GitHub stars, making it the fastest-adopted personal-agent software ever shipped. (One sovereignty caveat the lineage leaves open: it runs on your machine but typically calls a hosted frontier model for its tool loop — making the model itself local and sovereign is its own problem, since small models hold multi-step tool-use worst from prompt alone. That is the agentic-distillation work taken up under distil-then-deploy in → ch. 20, and the heavy-hands harness in → ch. 17.)
The architecture is the part to study, because it independently re-derives this chapter's whole history. An OpenClaw agent's identity is a workspace of plain markdown files, read into context at session start:
SOUL.md— persona, values, tone, behavioural limits (the character card, by another name)USER.md— who the human is (the entity-memory / predicate store)MEMORY.md+memory/YYYY-MM-DD.md— long-term memory plus daily working notes, which the agent edits itselfHEARTBEAT.md— a schedule of proactive, self-initiated activity (the Tamagotchi clock and the Ukagaka idle-talk timer, reborn)
Engine strictly separated from character; character as portable, human-readable, editable files; memory as documents the agent maintains; presence in the chat surfaces you already use; scheduled proactivity. That is ELIZA's script/engine split + AIML's portable persona + ChatScript's persistent user facts + Ukagaka's modular ghost + SmarterChild's live-where-the-user-is + Love Plus's real-clock continuity, in one stack — built by an assistant-tooling community that was, for the most part, not consciously drawing on any of it.
And the thirty-year pattern repeated on schedule: users immediately named their agents (the project's own mascot, Molty, set the tone), wrote them souls, and treated them as someones. Within weeks there was Moltbook, a social network populated by the agents themselves, and dedicated companion frameworks built on top (e.g. soulclaw: persona libraries, tiered memory, twenty-plus chat channels). Mainstream coverage oscillated between buzz and alarm — the alarm being legitimate: an always-on agent with credentials, tool access, and an open prompt surface inside your messaging apps is a genuinely new security exposure class (prompt injection with real hands; → ch. 22).
The lesson of the whole lineage, stated once: every sufficiently good assistant gets converted into a companion by its users — naming, persona, and affection are defaults, not niche behaviours. The conversion always happens to systems whose builders treated identity, memory, and proactivity as afterthoughts. The thesis of this book's runtime work (→ ch. 18) is simply to take the conversion as the design centre instead of the accident: build the always-on agent as a companion — soul, memory, heartbeat, and fiduciary duty first-class — rather than waiting for users to improvise one on top of an assistant.
→ For deeper per-product history see ch. 04.