Fine-tune an open model on your data, prove it beats the baseline, and serve it behind a stable API.
Recommended stack
- PyTorch + Hugging Face transformers/peft - training. LoRA fine-tunes run on one GPU and merge cleanly for serving.
- Weights & Biases (or MLflow) - experiment tracking. If it isn't logged, it didn't happen - configs, metrics, and artifacts per run.
- FastAPI + uvicorn - serving. Typed inference API with batching hooks and easy containerization.
- Docker - packaging. The training and serving environments must be reproducible to be debuggable.
Build steps
- Freeze an eval set FIRST (200+ labeled examples); score the base model - that's the bar.
- Prepare data into prompt/completion (or task-appropriate) pairs; version the dataset like code.
- LoRA fine-tune with small sweeps (lr, epochs); log every run with its config.
- Pick the run that beats baseline on the frozen eval - not on training loss.
- Merge adapters, containerize the FastAPI server, and load-test before swapping traffic.
Watch out for
- Evaluating on training data (or letting it leak into the eval set).
- Tuning on vibes instead of the frozen eval.
- Serving with a different tokenizer/version than training.
Definition of done
- Fine-tuned model beats baseline on the frozen eval set
- Any past run reproducible from its logged config
- P95 inference latency meets the product budget under load