KVM VPS from Rs 149/mo. Mumbai, Noida and Jaipur nodes.

24×7 infrastructure operations Sales +91 98297 14343
vpswala.in
VPSWala 3 min read

API Hosting on a Linux VPS: FastAPI, Worker Sizing & Process Supervision

A production architectural guide for hosting Python FastAPI backends on a Linux KVM VPS, covering asynchronous worker tuning, process supervision, and Nginx.

Building and deploying modern asynchronous application programming interfaces (APIs) using Python frameworks like FastAPI requires an infrastructure architecture that maximizes non-blocking I/O throughput. While prototyping on serverless platforms is common during early development, running production APIs on a dedicated Linux KVM virtual private server eliminates execution cold starts, unpredictable per-request pricing, and database connection pooling limitations.

The Asynchronous API Concurrency Model

FastAPI is built on Starlette and Pydantic, executing on top of an Asynchronous Server Gateway Interface (ASGI) runtime such as Uvicorn. Traditional WSGI servers allocate one thread or process per incoming HTTP request. In contrast, an ASGI event loop running on Python's asyncio allows a single process to handle thousands of concurrent client connections simultaneously, provided that downstream database and network queries utilize non-blocking asynchronous drivers (such as asyncpg or motor).

To operate a production-ready API on a Linux VPS, engineering teams structure the deployment into three coordinated layers:

  1. Worker Process Management: Gunicorn acting as a process manager utilizing Uvicorn's asynchronous worker class (uvicorn.workers.UvicornWorker).
  2. Service Supervision: systemd maintaining service uptime, managing environment variables, and capturing crash traces.
  3. Reverse Proxy: Nginx handling SSL termination, rate limiting, and request buffering before forwarding traffic to local application sockets.

Step 1: Worker Process Sizing and Gunicorn Configuration

Because Python's Global Interpreter Lock (GIL) prevents a single Python process from utilizing multiple physical CPU cores for compute-heavy tasks, running multiple worker processes is necessary to saturate multi-core KVM VPS instances.

Create a Gunicorn configuration file at /var/www/myapi/gunicorn_conf.py:

import multiprocessing

# Bind to internal loopback or UNIX socket
bind = "127.0.0.1:8000"

# Compute optimal worker count: (2 * cores) + 1
workers = (multiprocessing.cpu_count() * 2) + 1

# Specify Uvicorn's asynchronous worker class
worker_class = "uvicorn.workers.UvicornWorker"

# Maximum concurrent client connections per worker process
worker_connections = 1000

# Keepalive connection duration
keepalive = 5

# Graceful worker restart timeout
timeout = 30
graceful_timeout = 30

# Logging parameters
accesslog = "/var/log/myapi/access.log"
errorlog = "/var/log/myapi/error.log"
loglevel = "info"

Step 2: Configuring systemd Service Supervision

To ensure your API initializes on system boot and restarts automatically if an unhandled exception occurs, create a systemd unit at /etc/systemd/system/fastapi.service:

[Unit]
Description=Production FastAPI Backend Service
After=network.target

[Service]
User=apiuser
Group=apiuser
WorkingDirectory=/var/www/myapi
Environment="PATH=/var/www/myapi/venv/bin"
EnvironmentFile=/var/www/myapi/.env
ExecStart=/var/www/myapi/venv/bin/gunicorn -c /var/www/myapi/gunicorn_conf.py main:app

Restart=always
RestartSec=3

# Resource governance
LimitNOFILE=65536

[Install]
WantedBy=multi-user.target

Enable and start the service:

sudo systemctl daemon-reload
sudo systemctl enable --now fastapi.service
sudo systemctl status fastapi.service

Step 3: Configuring Nginx Reverse Proxy with Rate Limiting

Place Nginx in front of Uvicorn to protect the application from request surges and handle SSL termination:

# Rate limiting definition in http context
limit_req_zone $binary_remote_addr zone=api_ip_limit:10m rate=30r/s;

server {
    listen 80;
    server_name api.yourdomain.com;
    return 301 https://$host$request_uri;
}

server {
    listen 443 ssl http2;
    server_name api.yourdomain.com;

    ssl_certificate /etc/letsencrypt/live/api.yourdomain.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/api.yourdomain.com/privkey.pem;

    location / {
        limit_req zone=api_ip_limit burst=50 nodelay;

        proxy_pass http://127.0.0.1:8000;
        proxy_http_version 1.1;

        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;

        proxy_connect_timeout 60s;
        proxy_read_timeout 60s;
    }
}

Step 4: Database Connection Pooling and Graceful Lifespan Shutdown

In high-throughput asynchronous services, establishing a new database connection per HTTP request introduces latency and exhausts database connection pools. FastAPI supports the ASGI lifespan protocol to initialize persistent, shared database pools on application startup and gracefully close them when receiving a SIGTERM signal from systemd during rolling deployments.

When connecting to PostgreSQL using asyncpg or SQLAlchemy 2.0 with async engine drivers, configure connection pools inside the lifespan context manager:

from contextlib import asynccontextmanager
from fastapi import FastAPI
import asyncpg

# Lifespan context manager for resource management
@asynccontextmanager
async def lifespan(app: FastAPI):
    # Initialize connection pool on startup
    app.state.db_pool = await asyncpg.create_pool(
        dsn="postgresql://api_user:secret@127.0.0.1:5432/production_db",
        min_size=10,
        max_size=30,
        max_inactive_connection_lifetime=300.0,
        command_timeout=15.0
    )
    yield
    # Close pool cleanly on shutdown (SIGTERM / SIGINT)
    await app.state.db_pool.close()

app = FastAPI(lifespan=lifespan)

@app.get("/health")
async def health_check():
    async with app.state.db_pool.acquire() as conn:
        val = await conn.fetchval("SELECT 1")
    return {"status": "ok", "db": val}

Using connection pools ensures that active Uvicorn worker processes reuse open TCP sockets to PostgreSQL. When systemd stops or reloads the service, the lifespan context guarantees that in-flight SQL transactions complete before worker processes terminate, preventing data corruption and connection leaks.

Hardware Sizing for Production API Workloads

Asynchronous APIs require balanced CPU clock speeds and low-latency storage. While baseline REST services run smoothly on standard 2-vCPU KVM instances with 4 GB of RAM, high-concurrency workloads benefit substantially from high single-thread clock frequencies. VPSWala's AMD Ryzen 9 9950X VPS instances (featuring Zen 5 architecture, boost clock speeds up to 5.7 GHz, and DDR5 memory) deliver rapid JSON serialization, Pydantic data validation, and query dispatch.

Explore VPSWala Cloud VPS plans for versatile Linux hosting, or configure high-frequency compute on VPSWala 9950X VPS.


Sources

Not sure which size?

Send the stack, get a size.

Tell us the operating system, application stack, current traffic, database size and where it hurts today. You get a sizing recommendation, the matching plan and a price.

Related

More on this.

Get sizing help See VPS plans