field notes
Choosing a Vulkan allocator
Field note. Why the engine centralized its own allocator instead of adopting VMA. Only one of VMA's three benefits was actually about allocation, and that was the one with no evidence behind it.
Choosing a Vulkan allocator
Technical note: why the engine centralized its own allocator instead of adopting VMA.
Status: Accepted.
Landed in af254fd, before SSAO.
Context
SSAO is next on the roadmap.
It needs roughly four new render targets (linearized depth, an AO target, a blur ping-pong pair), and I opened VulkanImage.cs to start wiring them up.
Before writing anything new, I noticed I was about to produce the third copy of this function:
uint FindMemoryType(uint typeFilter, MemoryPropertyFlags properties):
memProps = queryPhysicalDeviceMemoryProperties()
for i in 0..memProps.MemoryTypeCount:
if (typeFilter has bit i)
and (memProps.MemoryTypes[i] has all `properties`):
return i
throw "no suitable memory type"
It already existed in VulkanBuffer.cs.
It already existed in VulkanImage.cs.
The ImGui backend had its own copy too.
Adding it to SSAO would have made four.
That’s the visible smell.
The deeper shape under it is worth describing, because it sets up why VMA was on the table at all.
The allocation path as it stood
Every buffer and every image in the engine went through the same five calls:
createBuffer(usage, data):
buffer = vkCreateBuffer(size, usage)
memReqs = vkGetBufferMemoryRequirements(buffer)
memory = vkAllocateMemory(memReqs.size, FindMemoryType(...))
vkBindBufferMemory(buffer, memory)
map → memcpy → unmap
return (buffer, memory)
The return type is a tuple, (buffer, memory), because the buffer and its backing memory are two separate Vulkan handles with two separate lifetimes.
Nothing in the type system enforced that they travel together; the discipline was entirely in the caller.
Destroy mirrored it: two calls out, in the right order.
Images added a third handle (the view), so their tuple was a three-tuple.
None of this is hard. It’s relentlessly mechanical, and the mechanical part was duplicated across every resource type.
The ImGui resize path
The place where raw allocation actually hurt was the ImGui backend. It kept per-frame vertex and index buffers that grew on demand, and the resize path looked like this:
if vtxSize > currentVertexBufferSize[frame]:
vkDestroyBuffer(vertexBuffer[frame])
vkFreeMemory(vertexMemory[frame])
(vertexBuffer[frame], vertexMemory[frame])
= createBuffer(VertexBuffer, vtxSize)
currentVertexBufferSize[frame] = vtxSize
Every time the frame’s geometry outgrew the current buffer, the entire backing DeviceMemory block was destroyed and a fresh one allocated.
vkAllocateMemory is not a cheap call.
Drivers cap the total number of live allocations (4096 on many implementations, lower on mobile), and we were spending them inside the frame loop.
This wasn’t broken.
ImGui’s geometry is tiny and resizes are rare.
But it was a preview: it’s the shape of code I’d keep writing as the engine grows.
The moment I start caching transient render targets or building a particle system, I’m back here writing destroy-then-reallocate against raw vkAllocateMemory.
Options
Three directions considered.
1. Stay raw, accept the duplication.
Copy FindMemoryType one more time for SSAO, keep the tuple lifetimes, keep the ImGui resize path.
Zero dependencies, zero new abstractions.
Costs a linear amount of boilerplate per new resource type and leaves the ImGui smell in place.
2. VMA (Vulkan Memory Allocator).
AMD’s GPUOpen allocator, via Silk.NET.Vulkan.Extensions.VMA on the C# side.
Three concrete changes:
- Sub-allocation.
A small number of large
vkAllocateMemorycalls feed many resources. The 4096-allocation ceiling stops mattering. ImGui’s resize path stops hitting the driver directly. - Memory type selection as intent.
You describe what the allocation is for (GPU-only, CPU-to-GPU upload, readback) and VMA picks the type.
FindMemoryTypedisappears, along with all three copies. - Coupled lifetime.
vmaCreateBufferreturns a buffer and a single opaqueVmaAllocation. No separateDeviceMemoryto track. Destroy isvmaDestroyBuffer(buffer, allocation). The tuple problem goes away at the type level.
3. One engine-owned allocator, no sub-allocation.
A single Allocator class that every resource path calls, holding the memory-type table once and returning a coupled (handle, Allocation).
Still exactly one vkAllocateMemory per resource.
This is the option the first draft of this note dismissed in a sentence, as “reinventing VMA poorly.” That dismissal was the mistake, and finding it is the reason the note is worth keeping.
Look again at VMA’s three benefits.
Only the first one is about allocation.
The other two are about where a piece of code lives and what type it returns, and neither needs a memory allocator to deliver.
Selecting a memory type by intent is a switch over two enum cases.
Coupling a lifetime is a readonly record struct returned alongside the handle.
Both are afternoon-sized, and both stop being a problem permanently once written.
Sub-allocation is the genuinely hard one, and it is the one with no evidence behind it. The engine allocates on the order of tens of resources: a handful of G-Buffer targets, a depth image, per-frame uniform and light buffers, a font atlas. Four new SSAO targets would not move that number near a driver limit. The 4096 ceiling was a real fact deployed in an argument about a codebase nowhere near it.
Decision
Build the engine-owned Allocator.
Defer VMA until something measurable asks for it.
Three copies of a helper function looked like an allocator problem. It was a code-organization problem wearing an allocator’s clothes, and the two came apart cleanly once I stopped treating VMA’s feature list as a single package.
What shipped:
Allocator.AllocateBuffer(state, size, usage, intent) -> (buffer, Allocation)
Allocator.AllocateImage(state, imageInfo, intent) -> (image, Allocation)
Allocator.DestroyBuffer(state, buffer, alloc)
Allocator.DestroyImage(state, image, alloc)
Allocator.Map(state, alloc) / Unmap(state, alloc)
MemoryIntent has two cases, GpuOnly and CpuToGpu, and PropsFor maps them to MemoryPropertyFlags.
Allocation is a readonly record struct (DeviceMemory, Size, MemoryType) that travels with the handle it backs.
FindMemoryType is one private method on Allocator, and the memory-type table is queried once at construction rather than per allocation.
VulkanBuffer and VulkanImage are now thin convenience wrappers over it.
The ImGui resize path got a separate fix that was never about the allocator: the per-frame buffers grow in doubling steps and stay mapped, so allocations fire O(log N) times during warm-up instead of on every resize. That is worth separating out. The expensive thing was reallocating on every growth, not the allocator underneath it. A growth policy fixed it, and VMA would have hidden it instead.
The weights:
- For: removes three duplicated helpers, collapses tuple lifetimes to a single handle, and makes SSAO’s four new render targets cost exactly the boilerplate they should cost. No new dependency, and no version lag between a C# binding and a native library.
- Against: the allocation-count ceiling is still real and still unaddressed. If a technique ever allocates per-frame or per-object at scale, this design walks into the wall VMA exists to prevent.
- The sharper against: this is a bet that the lab’s allocation volume stays low. It is a bet on the shape of a personal rendering lab, not on the shape of an engine, and it would be the wrong bet for anything shipping.
The cost structure is what decided it. The friction of the duplication was immediate and certain; the friction of the allocation ceiling is deferred and hypothetical. Paying for the certain one and leaving the hypothetical one measured-but-unbuilt is the trade the lab wants.
Consequences
Migration scope, as it landed.
VulkanBuffer.cs, VulkanImage.cs, and the ImGui backend.
Every vkAllocateMemory in the engine now routes through GpuState.Allocator.
(buffer, memory) and (image, memory, view) became (buffer, Allocation) and (image, Allocation, view).
VMA is still the upgrade path.
Nothing here forecloses it.
The engine now has exactly one place that calls vkAllocateMemory, which is the cheapest possible position from which to swap in a real allocator later.
The trigger is a measurement, not a feeling: a technique that pushes live allocations toward the driver’s limit, or a profile that shows allocation cost inside the frame.
This is recorded in the PRD’s constraints table so the trigger outlives the note.
Do not touch Handles.cs.
The opaque BufferHandle / ImageHandle pool is defined but unused.
Activating it is a separate refactor.
Bundling it with the allocator work was exactly the scope creep this note exists to avoid, and that held.
The techniques VMA would have fought. Worth recording, because they were the strongest argument for the option I did not take, and they are unchanged by the decision:
- Transient aliasing in a render graph: when the G-Buffer depth and a later blur target could physically share memory because their lifetimes don’t overlap. VMA supports aliasing, but it’s not the happy path; you end up managing custom pools and lifetimes yourself.
- Tile-based deferred on mobile: the optimal G-Buffer layout on tile GPUs is “lazily allocated, never leaves on-chip memory,” a
VK_EXT_subpass_shading-adjacent pattern that VMA has opinions about. - Sparse / residency-managed resources: not on the roadmap, but if they ever were, VMA would sit in the way rather than help.
Staying raw keeps all three of these open, which is a benefit I did not pay for and should not pretend I planned.
Functional angle
Worth naming because it’s the reason this decision was smaller than it looked.
The allocator lives entirely on one side of the functional core / imperative shell line.
vkAllocateMemory, the physical device queries, the memory-type table: all of that is in RenderLab.Gpu, the imperative shell.
The functional core, which describes frames as data (render graph nodes, pass descriptors, G-Buffer layouts), doesn’t know or care which allocator sits underneath.
That’s the test for whether an abstraction is in the right layer: can you swap it without the rest of the code noticing?
Here, yes, which is also why deferring VMA costs so little.
The decision was never load-bearing for anything above it.
Follow-ups (not in this note)
- Revisit VMA when a measurement, not an intuition, says allocation count or cost matters.
- Revisit when transient aliasing in the render graph becomes a real concern.
Handles.cspool activation: unblocked by this decision but not caused by it.