Run your own chatbot and API
Serve an open model behind a ChatGPT-style web interface with Ollama and Open WebUI, or behind an OpenAI-compatible API with vLLM so your own apps can call it.
It runs on a GPU you rent, not a shared public service, so your prompts are not used to train anyone else’s model.
Which GPU to pick
7B–8B models run fast on 16–24 GB. 14B–32B models need 24 GB with 4-bit weights, or 48 GB at higher precision. 70B-class models want 48 GB with 4-bit weights, or 80 GB and up.
Serve a chat model for an evening (A 13B model on vLLM, four hours online) costs, on today’s cheapest hosts:
| GPU | Typical time | Cost today |
|---|---|---|
| RTX 3090 | 4 h | USD 1.04 |
| RTX A5000 | 4 h | USD 1.48 |
| RTX 4090 | 4 h | USD 2.60 |
| RTX A6000 | 4 h | USD 2.60 |
| A100 SXM4 | 4 h | USD 3.12 |
Ready-made software
Pick one of these when you rent and it starts with everything installed:
- Ollama + Open WebUI (chat) — Run open models (Llama, Gemma, Qwen, Mistral…) with a ChatGPT-style web UI and the Ollama API on port 11434.
- vLLM (OpenAI-compatible API server) — High-throughput LLM serving with an OpenAI-compatible API on port 8000. Needs a 24 GB+ GPU.
- Text-generation WebUI (Oobabooga) — Chat, notebook and API modes for local LLMs, with model loaders for GGUF/GPTQ/EXL2.
How it works
- Create an account with Google or your email address. There is no ID upload: your first payment is how we know who you are.
- Add credits in your own currency.
- Add your SSH key, then pick the software and a GPU from the live board.
- Wait for Running — about a minute on a warm host, a few minutes on a cold one — then open the app or connect over SSH.
- Stop when you are done. GPU billing ends at once; only the disk is billed until you destroy the instance.
Questions
Is it compatible with the OpenAI API?
Yes, with vLLM. Point any OpenAI client library at your instance’s address and port and it works without code changes.
Can other people use my chatbot?
Yes, through the app link or the API port. Remember the instance is billed while it runs whether or not anyone is using it; the idle guard can stop it for you.
Which model should I start with?
Qwen, Llama or Gemma at 7B–8B are good first choices: fast, cheap to run, and strong at following instructions.