GPU requirements
Deploy Qwen2.5 7B (RTX 4090)
Deploy Qwen3 32B (A100)
Deploy Qwen3 235B MoE (2x H100)
Test the endpoint
Enable thinking mode
Add/think to the system prompt for step-by-step reasoning, or /no_think for direct answers:
Self-host Qwen3 or Qwen2.5 models with vLLM on a dedicated GPU. Covers model sizes from 7B to 235B.
/think to the system prompt for step-by-step reasoning, or /no_think for direct answers: