← library
prompt🧠 Machine learningv2 · updated 2026-06-12

Fine-tune and serve a model

PyTorch + Hugging Face + tracked experiments, served behind FastAPI.

Run it as a prompt

Paste this into any AI agent, or fetch it: curl -s https://uplift.page/api/v1/prompts/ml-fine-tune-serve/raw

prompt.md
# Fine-tune and serve a model

Fine-tune an open model on your data, prove it beats the baseline, and serve it behind a stable API.

## Recommended stack

- **PyTorch + Hugging Face transformers/peft** - training. LoRA fine-tunes run on one GPU and merge cleanly for serving.
- **Weights & Biases (or MLflow)** - experiment tracking. If it isn't logged, it didn't happen - configs, metrics, and artifacts per run.
- **FastAPI + uvicorn** - serving. Typed inference API with batching hooks and easy containerization.
- **Docker** - packaging. The training and serving environments must be reproducible to be debuggable.

## Build steps

1. Freeze an eval set FIRST (200+ labeled examples); score the base model - that's the bar.
2. Prepare data into prompt/completion (or task-appropriate) pairs; version the dataset like code.
3. LoRA fine-tune with small sweeps (lr, epochs); log every run with its config.
4. Pick the run that beats baseline on the frozen eval - not on training loss.
5. Merge adapters, containerize the FastAPI server, and load-test before swapping traffic.

## Watch out for

- Evaluating on training data (or letting it leak into the eval set).
- Tuning on vibes instead of the frozen eval.
- Serving with a different tokenizer/version than training.

## Definition of done

- Fine-tuned model beats baseline on the frozen eval set
- Any past run reproducible from its logged config
- P95 inference latency meets the product budget under load

The full prompt

Fine-tune an open model on your data, prove it beats the baseline, and serve it behind a stable API.

Recommended stack

  • PyTorch + Hugging Face transformers/peft - training. LoRA fine-tunes run on one GPU and merge cleanly for serving.
  • Weights & Biases (or MLflow) - experiment tracking. If it isn't logged, it didn't happen - configs, metrics, and artifacts per run.
  • FastAPI + uvicorn - serving. Typed inference API with batching hooks and easy containerization.
  • Docker - packaging. The training and serving environments must be reproducible to be debuggable.

Build steps

  1. Freeze an eval set FIRST (200+ labeled examples); score the base model - that's the bar.
  2. Prepare data into prompt/completion (or task-appropriate) pairs; version the dataset like code.
  3. LoRA fine-tune with small sweeps (lr, epochs); log every run with its config.
  4. Pick the run that beats baseline on the frozen eval - not on training loss.
  5. Merge adapters, containerize the FastAPI server, and load-test before swapping traffic.

Watch out for

  • Evaluating on training data (or letting it leak into the eval set).
  • Tuning on vibes instead of the frozen eval.
  • Serving with a different tokenizer/version than training.

Definition of done

  • Fine-tuned model beats baseline on the frozen eval set
  • Any past run reproducible from its logged config
  • P95 inference latency meets the product budget under load

Served from the uplift.page library and refreshed within 5 minutes of every update.