
AI inference can slow down even when a dedicated server still has plenty of CPU and memory available. On a multi-socket or multi-die system, the position of a process, its memory pages, network queues, and attached devices affects how quickly data moves through the request path. NUMA-aware CPU pinning gives operators a way to make that placement deliberate. It can reduce unnecessary thread migration and remote memory access, but it is not a magic switch. A reliable result comes from mapping the full workload, reserving capacity for the operating system, and measuring tail latency under a repeatable load.
Key Takeaways
- NUMA-aware placement keeps inference threads close to the memory and devices they use
- CPU pinning can reduce migration and contention, but only when housekeeping CPUs and failover capacity remain available
- NIC queues, storage, and accelerators must be mapped to the same locality as worker threads
- Measure p95 and p99 latency, first-token latency, throughput, remote-memory traffic, and CPU migrations before and after tuning
- Container platforms need CPU Manager, memory, and topology policies to preserve the placement you tested
- XLC dedicated servers provide a single-tenant base for repeatable CPU, memory, storage, and network tuning
What NUMA means on a dedicated server
Non-Uniform Memory Access, or NUMA, divides a multi-socket or multi-die server into locality domains. Each domain has CPUs and access to nearby memory, while it can also reach memory attached to another domain. The system remains one machine, but the cost of a memory access depends on where the data resides. This is why two servers with similar core counts can show different tail latency when an application repeatedly crosses NUMA boundaries.
On a bare metal server, the topology is visible and repeatable. Administrators can inspect it with lscpu, numactl -H, or hwloc, then decide how to place processes, memory, interrupts, storage queues, and devices. The goal is not to force every workload onto one node. The goal is to keep each latency-sensitive path local where that improves the measured result, while allowing less sensitive work to use the rest of the machine.
NUMA tuning is therefore a placement problem rather than a single switch. CPU affinity without memory affinity can leave threads running locally while reading remote pages. Memory binding without a matching network or accelerator path can move the bottleneck to I/O. A useful configuration begins with the complete request path and the server’s actual hardware topology.
This matters to XLC’s dedicated-server audience because a single-tenant machine removes a layer of unpredictable competition. Teams can benchmark the same CPU, memory, PCIe, storage, and network layout over time, then keep the settings as part of the deployment record.
Why AI inference exposes NUMA mistakes
Inference services often keep large model weights, token caches, vector indexes, request queues, and runtime buffers in memory at the same time. A request may also pass through a gateway, tokenizer, scheduler, model worker, and response stream. If those components are scattered across NUMA nodes, the service can spend more time moving data than executing useful work.
The first symptom is not always lower average throughput. Remote memory access, scheduler migration, cache misses, and queue handoffs can appear as higher first-token latency or a wider p95 and p99 distribution. A service may look healthy under a light test, then show unstable tail latency when concurrency increases and more workers compete for memory bandwidth.
Model-serving frameworks also use background threads for batching, tokenization, garbage collection, telemetry, and network processing. Pinning only the main worker does not guarantee that these threads remain close to the memory and devices they use. Measure the whole process tree and identify which threads are on the critical path before changing affinity.
A practical baseline should record the model, batch size, concurrency, request mix, sequence lengths, warm-up period, CPU frequency policy, memory usage, and software versions. Without those controls, a NUMA change can appear successful simply because the second test used a different workload or a warmer cache.
Tip: NUMA-aware tuning should be judged by the complete latency distribution and throughput under a repeatable request mix, not by one fast response from an idle server.
Map CPU, memory, network, and accelerator locality
Start by drawing the hardware map. Use lscpu and numactl -H to identify CPUs and memory nodes. Use lstopo or the device paths under /sys to see how NICs, NVMe devices, and accelerators attach to PCIe roots. The exact topology varies by server, so copy a real inventory into the runbook instead of relying on a generic socket diagram.
For a CPU-only inference service, place the worker threads and their model pages on the same NUMA node when the working set fits. For a multi-node model or a larger service, split workers into deliberate groups and give each group a clear memory policy. A balanced placement may be better than a single-node binding when the model or concurrency level exceeds one node’s memory bandwidth.
Network locality is part of the same decision. Multi-queue NICs can distribute receive processing across CPUs, but the selected queues and IRQs should be close to the application threads when the service is sensitive to packet-processing latency. RSS, RPS, RFS, and transmit steering can change where work is performed, so validate the queue map instead of assuming the default is optimal.
If an accelerator is present, align the CPU workers, host memory, and PCIe device with the device’s NUMA node where possible. Kubernetes Topology Manager and CPU Manager can help coordinate CPU and device placement for suitable pods, but the policies must match the resource requests and the cluster’s actual topology.
XLC dedicated servers can provide a stable single-tenant platform for this kind of mapping. Teams can benchmark the application path on a known hardware layout, keep the operating system and drivers controlled, and host the API, database, queue, monitoring, or model gateway layers alongside the inference service when that architecture fits.
The point is not to maximize affinity settings. It is to remove unnecessary crossings from the hot path, then leave enough unpinned capacity for system services, health checks, logging, updates, and recovery.
Tip: Record the NUMA node for every latency-sensitive CPU pool, NIC queue, NVMe device, and accelerator. If the map cannot be explained, the tuning is not ready for production.
Pin workers without starving the operating system
CPU pinning can reduce scheduler migration and keep a worker on predictable cores, but a server still needs housekeeping capacity. Reserve CPUs for the kernel, interrupts, network processing, storage work, monitoring, and orchestration. The Linux kernel documentation treats isolated CPUs and housekeeping CPUs as a trade-off: moving noise away from a workload also concentrates work elsewhere. Over-isolating a node can make the system less resilient and can increase latency when background tasks queue behind too few housekeeping cores.
Use the least restrictive control that produces a measured benefit. For a simple service, taskset, systemd CPUAffinity, a cpuset, or numactl may be enough. For a managed container platform, use CPU requests and limits, the Kubernetes CPU Manager static policy where exclusive CPUs are justified, and Topology Manager when CPU and device locality must be considered together.
Do not treat hyper-threading, core numbering, or CPU IDs as interchangeable. Full physical cores may behave differently from sibling hardware threads under heavy vector or memory workloads. Test the selected CPU set, document whether both siblings are assigned, and avoid placing unrelated latency-sensitive services on the same physical core.
Keep a fallback path. A pinned process can fail to start when the selected CPUs disappear, the server is resized, or a new node has a different topology. Use configuration validation at deployment time, expose the current CPU set in telemetry, and make rollback as easy as removing the affinity constraint.
Control memory placement and benchmark the result
Memory policy should follow the worker design. Linux supports policies such as local allocation, preferred nodes, binding, and interleaving, while cpusets can restrict the CPUs and memory nodes available to a task. Use numactl or libnuma deliberately, and remember that changing policy does not relocate pages that were already allocated. Start the process with the intended policy or use the application’s memory controls before the model is loaded.
Large model-serving workloads may benefit from huge pages or transparent huge page settings, but page size is not a substitute for locality. Test page faults, TLB behaviour, memory bandwidth, and allocation time with the actual runtime. Use numastat to compare local and remote allocations, and investigate numa_miss, numa_foreign, local_node, and other_node values instead of assuming that a high hit count alone proves the application is correctly placed.
Compare an untuned baseline with one controlled change at a time. Keep the model, runtime, request distribution, concurrency, warm-up, CPU governor, and background load consistent. Record p50, p95, and p99 latency, time to first token, tokens per second, request throughput, CPU migrations, context switches, local and remote memory traffic, memory bandwidth, NIC drops, and queue depth. Preserve the winning topology and tuning decisions in deployment configuration, then monitor them after kernel, driver, firmware, or model updates.
Conclusion
NUMA-aware CPU pinning can make bare metal AI inference more predictable by keeping workers, memory, network queues, and devices close to one another. The benefit depends on the workload and the topology. A strict binding that improves one benchmark can reduce capacity or resilience if it leaves too little room for the operating system, background work, or failover.
For teams building AI infrastructure today, XLC dedicated servers provide a practical single-tenant foundation for APIs, model gateways, databases, storage, and monitoring while the inference layer evolves. Measure the complete request path, keep the topology documented, and apply affinity only when the production workload shows a clear and repeatable improvement.


