Inside the LLM
A sentence that enters a language model does not immediately become an answer. The text passes through a series of processes within the model from tokenization and representation as embeddings, through processing with transformers and attention, to the prediction of the next token. Each stage plays a role in determining how the model processes context and produces its output.

From text to prediction
What sits around the model
The transformer is only one part of a language-model application. A chat system can add instructions, conversation context, retrieved documents, tools, and memory around it. The later chapters take a look at those pieces.
Foundation model
learns from large amounts of text
Instruction tuning
learns to follow instructions
Context
information from the conversation
RAG
documents retrieved when needed
Tools
actions outside the model
Memory
information kept across sessions
Agentic workflow
multiple steps and tool calls
Start with the text
Follow the process from the first token and see what happens along the way.