All courses

Mini-course · 7 questions

Preview

Post-training decisions: seven case studies

Stay with one question long enough to think for yourself.

Inside the course

A glimpse of the questions.

Question 01

Day 1: The support bot with yesterday's facts

Case: ParcelFlow's support assistant
All organizations, observations, and numbers in this case are fictional. Allow 10–15 minutes.

ParcelFlow sells shipping software. Its integration instructions change weekly. The assistant should answer from the current approved documentation, cite the relevant section, and ask for missing details instead of inventing an answer.

On 100 diagnostic tickets used for development, the current model gives 56 fully supported answers and follows the required response format 82 times. A clearer prompt with three good examples raises format compliance to 96 but fully supported answers only to 58. These checks can overlap. Reviewers find obsolete instructions and missing documentation in many remaining factual failures. You also have 8,000 historical support threads, some containing outdated advice and personal details.

Working vocabulary: Post-training changes a pretrained model through further training. Supervised fine-tuning (SFT) teaches it using examples of desired responses. Retrieval-augmented generation (RAG) supplies retrieved documents when the model answers; a basic RAG application can do this without changing model weights. An evaluation set is a collection of cases with a defined scoring rubric.

Your decision memo — at most 180 words:

  1. Choose the next experiment: more prompting, retrieval, SFT, or a combination. Explain which observed failure each part addresses.
  2. Describe the input and desired answer for one training example you would collect if a behavior gap remains. Decide what to do with the 8,000 old threads.
  3. Define two quality checks and a fair comparison on fresh cases. State what result would justify trying SFT after the first experiment.
Question 02

Day 2: Turn messy tickets into SFT demonstrations

Case: Alder's returns-triage assistant
All policies, data, and numbers below are fictional. Allow 10–15 minutes.

Alder wants a model to turn a customer message and a supplied policy excerpt into a structured recommendation. A person still approves the action. A good prompt produces valid structure, but reviewers see inconsistent decisions and guessed missing facts. You are planning an SFT pilot.

The supplied policy R7 says:

  • Proof of purchase is required. If its presence is unknown, ask for it before deciding.
  • Unopened items can be returned within 30 days.
  • Faulty items qualify for warranty review within 365 days.
  • If no rule applies or the policy is unavailable, send the case to a person.

The output must have exactly these fields: decision, policy_id, missing_fields, and reason. Allowed decisions are return, warranty, needs_info, and manual_review. Use an array for missing fields and a short explanation for the reason. Use purchase_proof for missing proof.

You have 2,400 historical ticket threads with repeat customers and copied messages. Some answers follow old policy R6. Another model generated 100 paraphrases from 20 of these tickets. A later month's untouched tickets are available for evaluation.

Working vocabulary: An SFT demonstration pairs the model's available input with a reviewed target response. Training data changes model weights; validation data helps choose a setup; a final test estimates performance after choices are fixed. Leakage occurs when evaluation information improperly influences training or model selection.

Your data-design memo — at most 200 words:

  1. Write the target object for each case: A: faulty item, 200 days old, proof present; B: unopened item, 12 days old, proof unknown. Both inputs include R7.
  2. Give three preparation rules for turning the historical threads into trustworthy demonstrations.
  3. Describe the training/validation/test split. Where do the 100 paraphrases belong, and what must be excluded from the model's input?
Question 03

Day 3: Adapt a small model within a compute budget

Case: FieldNote's technician assistant
All organizations, budgets, and dataset sizes are fictional. Allow 10–15 minutes.

FieldNote's assistant receives an equipment report and an approved manual excerpt. It must suggest one permitted next check, cite the section, and ask for a missing model number. A prompt-only small model still skips that last behavior.

The team has a compatible, licensed 3-billion-parameter model, 1,200 reviewed demonstrations, and separate validation and test sets. It can use one graphics processing unit (GPU) with 16 GB of memory for a short pilot. A colleague proposes: “Use QLoRA instead of SFT. The model will be 4-bit, so the whole training run needs only 1.5 GB.”

Working vocabulary:

  • SFT is the learning objective: make reviewed target responses more likely.
  • Full fine-tuning updates the model's ordinary weights.
  • LoRA, low-rank adaptation, freezes the original weights and trains small added parameter matrices called adapters.
  • QLoRA trains LoRA adapters while the frozen base weights use a memory-saving 4-bit representation. Four bits describe storage precision, not a four-choice answer space.
  • An epoch is one pass through the training examples; a checkpoint is a saved model state.

Your experiment memo — at most 180 words:

  1. Correct both parts of your colleague's claim. Calculate raw base-weight storage at 16 bits and 4 bits, using 8 bits per byte and decimal GB. Name two other things that occupy training memory.
  2. Choose an initial adaptation approach and describe what gets updated and what stays fixed.
  3. Propose a fair comparison against the existing model, including one small feasibility check, one behavior metric, and one regression check. Explain when you would stop or reject the adaptation.