Why Your AI Roleplay Character Forgets Everything (and How to Fix It)
Kryme
Your AI roleplay character forgets things because a language model only "sees" a limited amount of text at a time, called its context window. Every message, the character card and the instructions all have to fit inside it. Once a conversation grows past that limit, the oldest messages silently drop out of the prompt, and the character can no longer recall them.
This is the most common frustration in AI roleplay, whether you chat with a companion online or run your own setup with SillyTavern and a local model. The good news: it can be managed. In this guide you'll learn why memory loss happens, how to fix it in SillyTavern (summaries, lorebooks, vector storage), what each fix costs in hardware, and how to skip the setup entirely.
Why does my AI character forget things?
A large language model (LLM) has no memory between two replies. Each time you send a message, your app rebuilds a single prompt from scratch and sends it to the model. That prompt usually contains:
- The system prompt: general instructions on how to write.
- The character card: personality, appearance, backstory, scenario and example dialogue.
- The chat history: as many recent messages as will fit.
- Room for the reply: space reserved for the model's answer.
All of this is measured in tokens, small chunks of text. In English, one token is roughly three quarters of a word, so 1,000 tokens is about 750 words. The context window is the maximum number of tokens the model can read at once.
How many messages fit in a context window?
Here is a simple estimate. Assume a detailed character card of about 1,500 tokens, 500 tokens of instructions and reserved reply space, and an average exchange (your message plus the character's reply) of about 250 tokens.
| Context window | Space left for history | Approximate exchanges kept |
|---|---|---|
| 4k tokens | about 2,000 tokens | about 8 |
| 8k tokens | about 6,000 tokens | about 24 |
| 16k tokens | about 14,000 tokens | about 56 |
| 32k tokens | about 30,000 tokens | about 120 |
These are rough numbers: long, descriptive replies fill the window much faster. The key point is that everything beyond the limit is simply cut, usually starting with the oldest messages. The character is not ignoring you; the information is no longer in front of it.
Why a long character card makes it worse
The character card is sent with every single message. A 3,000-token card in an 8k window leaves less than half of the space for the actual conversation. Long example dialogues, detailed backstories and pasted lore all compete with your chat history for the same room.
Why does the character repeat itself?
Memory loss and repetition often go together. When the window only holds a handful of recent replies, the model sees its own phrasing over and over and tends to copy it. Repetition penalties help, but the root cause is often a crowded context.
How to fix memory in SillyTavern
SillyTavern is a popular open-source frontend for AI roleplay. It does not run a model itself: it connects to a backend such as KoboldCpp, Ollama or LM Studio (or to a cloud API). Here are the main tools it offers against forgetting, from the simplest to the most advanced.
1. Increase the context size
The most direct fix is a bigger window. You need to raise it in two places: in your backend (the context length the model is loaded with) and in SillyTavern's context size setting. If the two do not match, either the extra space is wasted or the prompt gets truncated by the backend. Some backends start with a fairly small default context, so check it before anything else.
A bigger context costs video memory (VRAM) and makes each reply slower to start, especially on long chats. See the table in the next section.
2. Use the Summarize extension
Summarize is an extension installed with SillyTavern by default, available in the Extensions panel. It periodically asks a model to write a summary of what has happened in the story and inserts that summary into the prompt, so older events survive after the original messages drop out.
The SillyTavern documentation itself suggests treating it as long-term memory "with a grain of salt": summaries are written by a language model, so they can lose details or invent things. You can read and edit the current summary, roll it back to a previous one, or pause automatic updates, and it is worth checking it regularly on long stories.
3. Write lorebooks with World Info
World Info, also called lorebooks, works like a dictionary of facts. Each entry has activation keywords and a description. When one of the keywords appears in the recent conversation, SillyTavern inserts that entry into the prompt.
Lorebooks are ideal for stable facts: the name of a city, a character's family, the rules of a fictional world, or an important event you want the character to remember. They only cost context space when they are relevant. The downside is that you write and maintain the entries yourself.
4. Turn on Vector Storage
Vector Storage can be enabled for chat messages (the documentation calls this chat vectorization). It turns each message into a vector, a list of numbers that represents its meaning. When you write a new message, SillyTavern searches the whole chat for older messages that seem related and slips the most relevant ones back into the prompt, even if they are far beyond the context window.
This is a form of retrieval-augmented generation (RAG). It needs a vectorization model to be configured, and the documentation is clear that it does not guarantee better memory: it retrieves messages that look similar, which is not always what the story needs.
5. Use the Author's Note for what matters right now
The Author's Note, found in the Options menu next to the chat input, inserts a short text into the prompt at a position and frequency you choose. Placed close to the end of the chat history, it has a strong influence on the next reply. It is a simple way to pin a few current facts, for example where the scene takes place or what the characters just agreed on.
6. Keep the character card short
Every token in the card is a token taken from your history. Trim repeated adjectives, move background lore into a lorebook, and keep example dialogue to a few strong lines. A lean card often improves memory more than any extension.
Context size vs VRAM
When you run a model locally, its context needs memory on top of the model itself (the "KV cache"). The figures below are rough orders of magnitude for a model quantized to about 4 bits, with an unquantized cache. Real numbers vary with the model architecture, the backend and its settings; quantizing the cache can roughly halve the context cost.
| Context size | Typical 7B to 8B model (approx. total VRAM) | Typical 12B to 14B model (approx. total VRAM) |
|---|---|---|
| 4k tokens | about 5 to 6 GB | about 8 to 9 GB |
| 8k tokens | about 6 GB | about 9 to 10 GB |
| 16k tokens | about 7 GB | about 10 to 12 GB |
| 32k tokens | about 9 GB | about 13 to 15 GB |
| 64k tokens | about 13 GB | about 18 to 22 GB |
In practice, an 8 GB graphics card is comfortable with a small model at 8k to 16k tokens, while 32k tokens and beyond on a mid-size model calls for 16 to 24 GB of VRAM. If you are curious about running AI tools on your own machine in general, our guide to ComfyUI covers the same hardware trade-offs for image generation.
Why it's still a hassle locally
All of these fixes work, but together they turn a chat into a small engineering project:
- Two programs to maintain. A backend to load the model and a frontend to chat, each with its own updates and settings.
- Settings that must agree. Context size, prompt format, sampler settings and stop sequences all have to match the model you picked.
- VRAM limits. More memory means more context, and more context means a more expensive graphics card or slower replies.
- Summaries to babysit. Automatic summaries drift, drop details or invent them, and need regular checking on long stories.
- Lorebooks to write. World Info is powerful, but every entry is written and updated by hand.
- Vector search to tune. Retrieval needs an extra model and still pulls in irrelevant messages from time to time.
If you enjoy tinkering, this is part of the fun. If you just want a character that keeps up with your story, there is a simpler path.
Skip the setup: AI companions on Yamete.gg
Yamete.gg handles conversation memory for you, with nothing to install or configure. Instead of asking you to manage summaries and lorebooks, the platform keeps track of your story in the background while you chat:
- The story is summarized as it unfolds, so earlier events keep shaping the conversation after the original messages are long gone.
- Key facts are remembered, such as your name and the details you share about yourself.
- The relationship and the scene are tracked: how close you and the character have become, where you are, and what is happening right now.
- The scenario evolves with your story instead of resetting to its starting point.
- Repetition is detected and steered away from, and characters are guided not to take over your side of the story.
- No GPU and no setup: it runs in your browser, on desktop or mobile.
SillyTavern vs Yamete.gg: memory out of the box
| SillyTavern (local setup) | Yamete.gg | |
|---|---|---|
| Setup | Backend, model, frontend and extensions to install and configure | None, open the site and chat |
| Story summary | Optional extension you configure, check and correct | Automatic, in the background |
| Facts about you | Written by hand in your persona or a lorebook | Picked up from the conversation automatically |
| Relationship and scene tracking | Not built in, notes kept by hand | Tracked automatically at every message |
| Evolving scenario | The card's scenario stays as written | Updated as the story moves on |
| Repetition | Sampler penalties to tune | Detected and steered away from automatically |
| Hardware | A GPU with enough VRAM for the context you want | Any browser, desktop or mobile |
To be clear about the limits: memory is kept within each conversation, and it is a condensed memory, not a word-for-word transcript, so very small details from long ago can still fade. For most stories, it means you can come back to a chat days later and pick up where you left off.
Try Yamete.gg and meet a companion that keeps up with you.
Yamete.gg is an adults-only platform (18+).
AI Roleplay Memory FAQ
Why does my AI character forget things?
Because a language model can only read a limited number of tokens at once, its context window. When the conversation grows past that limit, the oldest messages are removed from the prompt and the character can no longer see them.
How much context do I need for roleplay?
8k tokens is a workable minimum for short scenes, 16k tokens is comfortable for most stories, and 32k tokens or more helps with long campaigns. Bigger windows need more VRAM and make replies slower to start.
Does SillyTavern have long-term memory?
Not by default in the sense of remembering everything. It offers tools that act as long-term memory: the Summarize extension, World Info lorebooks and Vector Storage for chat messages. Each one needs to be enabled and checked by you.
What is the difference between a lorebook and a summary?
A summary is written automatically from what happened in the chat and changes as the story moves on. A lorebook is a set of entries you write yourself, inserted only when their keywords come up. Summaries track events, lorebooks hold stable facts.
Why does my AI repeat itself?
When the context only holds a few recent replies, the model keeps seeing its own phrasing and tends to reuse it. A shorter character card, a larger context and moderate repetition penalties usually help.
Do online AI companions remember past conversations?
It depends on the service. Some keep memory only within a single chat, others carry a few profile details across chats. On Yamete.gg, memory is kept within each conversation, and the story is summarized automatically so you can return to it later.