Our services

If you can think it, we can make it brainsoft.

Rate limiting and caching at the reverse-proxy layer

Written By: BrainSoft In DevOps

When you put a reverse proxy in front of your application, you gain a choke point where you can control traffic. Two of the most useful things to do there are rate limiting and caching. Without them, your origin servers are exposed to abuse and redundant load.

In this article I'll walk you through setting up both in Nginx, the most common reverse proxy. You'll define shared memory zones for rate limits, write rules to apply them selectively, and configure caching with proper cache keys and invalidation. You'll also see how to test your setup and avoid common pitfalls.

Setting up rate limiting

Rate limiting in Nginx uses the limit_req_zone directive to define a shared memory zone that tracks request rates per key (typically the client IP). Then you apply limit_req in the location where you want enforcement.

  1. Define the zone in the http block. For example, allow 10 requests per second per IP with burst of 20:
http {
    limit_req_zone $binary_remote_addr zone=api_limit:10m rate=10r/s;
}
  1. Apply the limit in a location block. Use burst to allow a queue, and nodelay to process queued requests immediately up to the burst limit:
location /api/ {
    limit_req zone=api_limit burst=20 nodelay;
    proxy_pass http://backend;
}
  1. Return a custom response when the limit is exceeded. By default Nginx returns 503. For APIs, a 429 is more accurate. Use limit_req_status:
limit_req_status 429;

You can also set a limit_req_log_level to see rejections in logs.

Configuring caching

Caching at the reverse proxy reduces load on your origin by serving repeated requests from a fast local store. Nginx caching is simple and effective, but you need to define a cache path, a key, and how long to keep entries.

  1. Define the cache zone. In the http block, set a path and shared memory zone:
http {
    proxy_cache_path /var/cache/nginx levels=1:2 keys_zone=my_cache:10m max_size=1g inactive=60m use_temp_path=off;
}
  1. Enable caching in a location. Set the cache zone, cache key, and bypass conditions. A typical setup for GET requests:
location / {
    proxy_cache my_cache;
    proxy_cache_key "$scheme$request_method$host$request_uri";
    proxy_cache_valid 200 60m;
    proxy_cache_valid 404 1m;
    proxy_cache_bypass $http_cache_control;
    proxy_pass http://backend;
}
  1. Add cache control headers. You can pass the upstream's cache headers and add X-Proxy-Cache to see if a request was a hit:
add_header X-Proxy-Cache $upstream_cache_status;
proxy_ignore_headers Cache-Control Expires;
proxy_hide_header Set-Cookie;

Be careful with dynamic content. Cache only what you know is safe, and use proxy_cache_bypass for requests with cookies or authorization headers.

Combining both and testing

You can apply rate limiting and caching in the same location. A typical pattern is to rate limit before checking the cache, so that even cached responses are throttled. Nginx processes directives in order, so put limit_req before proxy_cache in the location block.

location /api/ {
    limit_req zone=api_limit burst=20 nodelay;
    proxy_cache my_cache;
    proxy_cache_key "$scheme$request_method$host$request_uri";
    proxy_cache_valid 200 60m;
    proxy_pass http://backend;
}

After changing the configuration, test with nginx -t and reload. To see if caching works, use curl -I and check the X-Proxy-Cache header. For rate limiting, you can send a burst of requests with ab or a simple script and watch for 429 statuses.

When you need to invalidate the cache after a deploy, you can use the proxy_cache_purge directive from the ngx_cache_purge module or simply delete the cache directory. If you're managing many services, you might prefer a dedicated CDN or API gateway, but Nginx gives you a solid foundation. If you need help designing your architecture, get in touch.

Frequently asked questions

What is the difference between rate limiting and throttling?

Rate limiting restricts the number of requests a client can make in a given time window, often rejecting excess requests. Throttling usually means slowing down requests or allowing a burst but with a queued delay. In Nginx, limit_req can be configured to do either by using the burst parameter alone (throttling) or with nodelay (rate limiting with immediate rejection of overflow).

How do I decide what to cache at the reverse proxy?

Cache responses that are identical for many users, such as public static assets, API responses without user-specific data, or pages that change infrequently. Avoid caching responses with Set-Cookie headers or that vary by Authorization headers. Use proxy_cache_bypass for those requests.

Can I use rate limiting and caching with WebSocket connections?

WebSocket connections are long-lived and not typically cached. Rate limiting can still apply to the initial handshake request, but once the connection is upgraded, the proxy passes data through without further rate checks. For WebSocket support, you need to configure the upgrade headers in Nginx.


#DevOps