HugePages for PostgreSQL on Linux

HugePages memory pages are locked in system memory and have a much bigger page size than the default, so it benefits software programs that demand a large amount of memory such as a database. Using HugePages reduces sys (kernel) CPU usage thus improves system performance. The following is the procedure to set it up for PostgreSQL on Linux.

If a server is dedicated to one PostgreSQL instance, as a starting point, we can dedicate a quarter of the system memory to the shared buffers of the instance as recommended in the PostgreSQL documentation, e.g. 8 GB on a server with 32 GB memory. If your application never runs a SQL that requires a huge amount of memory, you sure can increase shared_buffers but note that "it is unlikely that an allocation of more than 40% of RAM to shared_buffers will work better than a smaller amount", says the documentation.

Suppose you can have a short maintenance window. After you set shared_buffers (to '8GB'), max_connections, and possibly wal_buffers in postgresql.conf, bring down the instance, and run postgres -D $PGDATA -C shared_memory_size_in_huge_pages. You may get a number such as 4218, which is the Huge Pages the instance requires based on a calculation according to your settings. As root, let's try to allocate that many Huge Pages on the fly. (Some recommend one more page as a guard page, i.e. 4219 in this case, which is fine.)

# echo 4218 > /proc/sys/vm/nr_hugepages; cat /proc/sys/vm/nr_hugepages

If cat shows less than you want, 4218 here, try a few more times. If it's still less, you have to reboot because there's not enough contiguous free memory. But put vm.nr_hugepages = 4218 in /etc/sysctl.conf, and also run grubby --args="transparent_hugepage=never" --update-kernel ALL before you reboot. (I'll explain this later.)

Then start the instance. Let's check HugePages at the OS level

# grep Huge /proc/meminfo
AnonHugePages:         0 kB
ShmemHugePages:        0 kB
FileHugePages:         0 kB
HugePages_Total:    4218
HugePages_Free:     4117
HugePages_Rsvd:     4117
HugePages_Surp:        0
Hugepagesize:       2048 kB
Hugetlb:         8638464 kB
The above means that the system has allocated 4218 HugePages memory pages, 4218-4117=101 pages really used, 4117 pages "free" (I'd rather call it available) and reserved, and 4117-4117=0 pages wasted. HugePages_Free will always be bigger than or equal to HugePages_Rsvd and the difference is wastage, i.e. allocated but never to be used. If you had allocated 4219 earlier, you would see a difference of 1 between them.[note1]

Inside psql, you should see

postgres=# select name, setting from pg_settings where name like '%huge%';
               name               | setting
----------------------------------+---------
 huge_page_size                   | 0
 huge_pages                       | try
 huge_pages_status                | on
 shared_memory_size_in_huge_pages | 4218
where 4218 is the amount Postgres calculated when we ran the postgres -D $PGDATA -C shared_memory_size_in_huge_pages command earlier offline.

To use the value in the future across reboots, remember to update /etc/sysctl.conf with vm.nr_hugepages = 4218.

Earlier, you may have run the grubby --args="transparent_hugepage=never" --update-kernel ALL command (from Red Hat). That disables transparent HugePages, which would cause sporadic CPU spikes, non-contiguous memory chunks, and other problems. It's strongly recommended we disable it. If you have not done so, run it and reboot the server when you have a chance.[note2]

If you plan to change shared_buffers, max_connections, or wal_buffers in PostgreSQL again, you need to shut down the instance and check the calculated HugePages again and set HugePages at the OS level accordingly. You can do a rough estimate first by adding shared_buffers and wal_buffers and 40 KB times max_connections (note the units they use!), plus a little more (locks e.g., and other little things as in pg_shmem_allocations). But to actually make changes, instance downtime is still needed.

If you want to guarantee all shared memory is in HugePages and not regular pagesize memory, set huge_pages to 'on' (equivalent to the value 'only' of the use_large_pages parameter in Oracle), so the instance will refuse to start when the requirement is not met.

If you really want to see HugePages memory being used by the instance, run view /proc/$(pgrep -f "^postgres: checkpointer")/smaps on OS, and search for KernelPageSize:     2048 kB (note the 5 white spaces) or MMUPageSize:        2048 kB. In our case, we see this memory segment with a Size of 8638464 kB, which is exactly 8436 MB, or 4218 HugePages pages of 2 MB pagesize. (Ignore the name /anon_hugepage (deleted). That's just a sign of POSIX mmap type shared memory allocation[note3], because this memory region is not backed by any file.)

Lastly, since you have configured HugePages, you may as well tune a few VM parameters as well. For example, set vm.overcommit_memory to 2 according to the advice of documentation (Enterprse DB recommends 0), vm.swappiness to 10 according to Enterprise DB (Red Hat used to recommend it too), etc.


Although not as strongly as for Oracle, HugePages has been recommended for PostgreSQL. It will become even more important once the Linux kernel is version 7 according to one study. See my brief summary in plain language.

__________________________

[note1]
To understand the 3 lines of HugePages_* in /proc/meminfo, look at this simple diagram

UUUUUFFFF <-- Total split into really used (U) and free (F)
UUUUURRR. <-- Total split into really used (U), reserved (R) and really free (.)
If one letter or dot is one HugePage, the above says
HugePages_Total: 9
HugePages_Free:  4
HugePages_Rsvd:  3
and you’ll have 4–3=1 page completely wasted.

[note2] Before and after reboot, you can check transparent HugePages by
cat /sys/kernel/mm/transparent_hugepage/enabled #should have [always] before reboot, [never] after
grep AnonHugePages /proc/meminfo #should have a non-zero value before reboot, 0 after
cat /proc/cmdline #should have transparent_hugepage=never appended after reboot
If the above grubby command doesn't work, see Section 7 of the article on Oracle HugePages for older Linux versions.

[note3]
Speaking of POSIX mmap type shared memory, if for some obscure reason you decide to use the traditional SysV mechanism to allocate HugePages (shared_memory_type set to 'sysv' instead of the default 'mmap'), you need to do more work. For example, postgres soft memlock limit and postgres hard memlock limit should be added to /etc/security/limits.conf, where limit can just be the system memory in KB to make it simple (it's meaningless but harmless to set it bigger than that). In /etc/sysctl.conf, kernel.shmmax should be set to a large number, which can be system memory in bytes (note: not KB!). session required pam_limits.so should be added to /etc/pam.d/login.


August - October 2026
(originally posted on LinkedIn but completely rewritten here)

Contact me
To my Computer Page