comez.dev
en
~/article/php-fpm-vs-php-cgiAll articles

article2 min read25 views

PHP-FPM vs PHP-CGI: what changes in production

Why process management matters more than the PHP version when a site starts returning 502 and 504 errors.

On this page (5)
  1. The difference in one paragraph
  2. Pools are the real feature
  3. Sizing pm.max_children
  4. Reading the symptoms
  5. Checklist

Most slow or unstable PHP sites I am asked to look at have nothing wrong with their code. The problem is how PHP processes are started, shared and limited.

The difference in one paragraph

With classic CGI, the web server starts a new PHP process for every request and throws it away afterwards. That is simple and safe, and slow. PHP-FPM (FastCGI Process Manager) keeps a pool of long-lived workers that handle many requests, which allows opcode caching, persistent connections and predictable memory use.

Pools are the real feature

Each site can have its own pool with its own user, limits and settings.

; /etc/php/8.3/fpm/pool.d/example.conf
[example]
user = example
group = example
listen = /run/php/example.sock
listen.owner = www-data
listen.group = www-data

pm = dynamic
pm.max_children = 20
pm.start_servers = 4
pm.min_spare_servers = 2
pm.max_spare_servers = 6
pm.max_requests = 500

request_slowlog_timeout = 5s
slowlog = /var/log/php-fpm/example-slow.log

Sizing pm.max_children

Do not guess. Measure the average memory of a worker under real load and divide the memory you can spare by that number.

ps --no-headers -o rss -C php-fpm8.3 | awk '{s+=$1; n++} END {print s/n/1024 " MB avg"}'

If the result is 60 MB and you can spare 2 GB, the ceiling is roughly 30 workers. Going above it trades slow requests for swapping, which is worse.

Reading the symptoms

  • 502 Bad Gateway usually means the pool crashed, the socket is missing, or permissions on the socket are wrong.
  • 504 Gateway Timeout usually means every worker is busy. Check pm.max_children and the slow log before raising timeouts.
  • Memory creeping up is often fixed by a sensible pm.max_requests, which recycles workers.

Checklist

  1. One pool and one system user per site.
  2. OPcache enabled and sized for the code base.
  3. pm.max_children derived from measured memory.
  4. The slow log switched on, and actually read.

General

Comments0

No comments yet. Be the first to comment.

Leave a comment

Related

Working on something similar?
Happy to review an architecture or help with delivery.

contact