Fine-Tune or Self-Host?
With the LLM foundations in hand, the real decision is which model — or which mix of models — to use, and whether you ever need to fine-tune or self-host at all.
Match the Model to the Task
Balance four things against what the task actually needs. A smaller model is often enough — and cheaper and faster.
- Capability — is it strong enough for the hardest step?
- Cost — what does each call add up to at scale?
- Latency — how fast must the answer come back?
- Context size — how much must fit in one prompt?
Use a Mix of Models
You rarely need your strongest model everywhere. Route a cheap, fast model to the easy steps and reserve a stronger one only where the task demands it.
When to Fine-Tune
Fine-tuning teaches a model your specific task or style from your own examples — but it is not the first move.
- Try prompting first — clear instructions and examples go a long way
- Then grounding — give it your data at answer time
- Fine-tune only when those still fall short and you have good, consistent data
When to Self-Host
Self-hosting means running an open model on your own infrastructure — for control, privacy, or cost at scale. It trades the convenience of a hosted API for real operational burden.
Build It
How to implement: pick a default model for your capstone and write one line on why — capability vs cost vs latency. Then note what would make you switch.
- Weekly AI Tasks tracker — start with a hosted model via API; consider a smaller, cheaper one for the frequent parse step; fine-tune only if prompting stalls.
- Personal brand site — a general hosted model is plenty; no fine-tuning or self-hosting needed.
What you learned
Module 5 gave you the LLM engineering decisions: choose the right model, mix models for cost and quality, and reach for fine-tuning or self-hosting only when prompting and grounding are not enough.