1. Prepare your input file
Create a JSONL file with one prompt per line:2. Deploy and upload
3. Install vLLM and upload the batch script
4. Download results and tear down
Monitor progress
Cost estimate
Tips
- vLLM offline batch mode uses continuous batching — much faster than sequential API calls.
- For 100K+ prompts, split the file and process in chunks to avoid OOM.
- Check your balance before starting:
runcrate billing balance.