Back to Generative AI Notes
Topic #16

How LLMs Work

An LLM works in two distinct phases: an expensive, one-time training phase that produces the model's weights, and a fast, repeated inference phase that uses those fixed weights to generate responses to new prompts.

Training vs Inference — The Distinction That Matters Most

TrainingInference
When it happensOnce (or periodically, for model updates), before the model is releasedEvery time someone sends a prompt
What it doesAdjusts millions/billions of weights to minimize prediction errorUses the already-fixed weights to predict the next token, repeatedly
CostExtremely high — large compute clusters over weeksMuch lower per request, but scales with usage volume
Who does itThe model provider (or you, if fine-tuning)Anyone calling the model via an API or running it locally

When you call an LLM API, you are only ever doing inference — the weights don't change based on your conversation. See LLM Inference for the mechanics of that process.

Training Has Its Own Stages

Raw text (internet, books, code)
  ↓ pretraining
Base model (fluent, but not great at following instructions)
  ↓ instruction tuning
Instruction-following model
  ↓ alignment
Aligned, deployable model (what you interact with via chat)

Each stage is covered separately: Pretraining, Instruction Tuning, Alignment.

Where Prompting and Fine-Tuning Fit In

ApproachChanges the model's weights?When to use
PromptingNoDefault choice — flexible, no training cost, works immediately
RAGNo — adds external context at inference timeWhen answers need to be grounded in specific, current, or private data
Fine-tuningYes — further trains weights on your own dataWhen you need consistent behavior/format/style that prompting alone can't reliably achieve — see Fine-Tuning vs Prompting

Common Mistakes

  • Assuming a conversation "teaches" the model anything permanent — nothing about an API conversation updates the underlying weights; the illusion of learning within a session comes entirely from context being re-sent each turn
  • Reaching for fine-tuning to add knowledge the model doesn't have — that's what RAG is for; fine-tuning is much better suited to changing behavior/style than injecting facts

Interview Relevance

"Does chatting with an LLM change the model?" is a deceptively simple, very common screening question — the correct answer is no, and being able to explain the training/inference split clearly is a good signal.

Practice Question

A user says "I told the chatbot my name yesterday, why doesn't it remember today in a new conversation?" Explain the answer using the training/inference distinction.

Want to go beyond the notes?

Join Coding Now Tech Institute's Generative AI course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available

How LLMs Work – FAQs

Quick answers about learning How LLMs Work in Generative AI.

This free note from Coding Now Tech Institute explains How LLMs Work in Generative AI — concept, syntax and worked code examples you can copy, run and revise before interviews.
Yes. Every Generative AI topic on Coding Now Tech Institute, including How LLMs Work, is 100% free with no signup required.
With focused practice, most students grasp How LLMs Work in 1–3 days from these notes; pairing it with Coding Now Tech Institute's mentor-led course takes you to job-ready depth faster.
Use the code examples in this note, then ask doubts for free on the Coding Now Tech Institute Community (/community) — expert instructors answer within 24 hours.
Call NowEnroll Now