04 foundation
The frame composes itself
The lighting pass, the tonemap pass, and the render graph. Composing passes as pure data, compiling order and barriers from declarations, and testing a frame without a GPU. Closes the introduction.
The frame composes itself
A render graph does not execute passes. It describes which passes could exist, and a pure function works out the rest.
TL;DR
- The lighting pass draws one fullscreen triangle and reads the G-Buffer through three samplers. It knows nothing about the geometry that produced them.
- The tonemap pass maps HDR into a displayable range, and deliberately does no gamma correction, because the target’s sRGB format does that in hardware.
- Three passes sharing resources create three problems that one pass never has: execution order, synchronization, and image layout transitions.
- Passes are declared as immutable records of what they read and write. A pure function topologically sorts them and derives every barrier from the usage changes.
- That function is 100 lines, has no idea Vulkan exists, and is covered by 18 tests that run without a GPU.
- All the Vulkan lives in one executor: barriers, then a recorder per pass.
- This is the last post of the introduction. It ends with a complete frame and no cliffhanger.
Picking up where the light went out
The last post ended in the dark. We built a G-Buffer: three textures holding everything a lighting algorithm needs, written by a geometry pass that refuses to compute a single light. Every visible pixel has a position, a direction, and a color. The data is complete.
But the screen shows nothing. The G-Buffer is a return value with no caller. No pass reads it, no pass maps it to something a display can show, and nothing connects the passes that would make the frame work. We have decomposition without composition.
This post fixes that, and then it stops. We will write the lighting pass, add a tonemap pass, and face the part that turns out to be more interesting than either of them: composing multiple passes into a correct, synchronized execution order without hand-writing that order anywhere. The answer is a render graph, and in this engine it is built entirely from immutable data and one pure function.
It is also the last post in this introduction. Four posts ago there was a motivation and no code. By the end of this one there is a window, a device, a swapchain, a pipeline, a G-Buffer, a lit image, and a system that composes passes into a frame, which is a complete Vulkan renderer in the sense that matters at this stage: nothing structural is missing from the path between a scene and a pixel.
The lighting pass has no geometry
The lighting pass does not draw the scene. It has no vertex buffer, no index buffer, and no mesh data, because the scene has already been drawn: it is sitting in three textures. What the pass draws is one triangle big enough to cover the screen.
That triangle needs no vertex buffer either. It comes out of the vertex index, which is the whole shader:
// fullscreen.vert, in full.
#version 450
layout(location = 0) out vec2 uv;
void main() {
// Fullscreen triangle from vertex index (3 vertices, no vertex buffer)
uv = vec2((gl_VertexIndex << 1) & 2, gl_VertexIndex & 2);
gl_Position = vec4(uv * 2.0 - 1.0, 0.0, 1.0);
}
Three vertices at (0,0), (2,0) and (0,2) in UV space, which lands as a triangle whose corners sit off-screen on two sides. The rasterizer clips it, and every pixel gets covered exactly once. A fullscreen quad made of two triangles would also work and would run marginally slower, because the diagonal seam makes the GPU shade the quads along it twice. Every fullscreen pass in this engine shares this one vertex shader: same three vertices, different fragment shader.
The fragment shader is where the frame gets its light. It reads the G-Buffer through three samplers bound as a single descriptor set, and for each pixel it reconstructs the surface the geometry pass saw:
// lighting.frag, reduced to the shape this post is about.
// The real file has grown a lot; the last section covers what it grew into.
layout(set = 0, binding = 0) uniform sampler2D gPosition;
layout(set = 0, binding = 1) uniform sampler2D gNormal;
layout(set = 0, binding = 2) uniform sampler2D gAlbedo;
layout(location = 0) in vec2 uv;
layout(location = 0) out vec4 outColor;
void main() {
vec3 fragPos = texture(gPosition, uv).rgb;
vec3 N = normalize(texture(gNormal, uv).rgb);
vec3 albedo = texture(gAlbedo, uv).rgb;
// One point light, in the units the scene stores it in.
vec3 lightDir = normalize(lightPos - fragPos);
float dist = length(lightPos - fragPos);
float atten = 1.0 / (1.0 + 0.09 * dist + 0.032 * dist * dist);
vec3 direct = max(dot(N, lightDir), 0.0) * albedo * lightColor * intensity * atten;
// Hemispheric ambient: ground color to sky color along the normal's up axis.
float skyMix = N.y * 0.5 + 0.5;
vec3 ambient = mix(ambientGround, ambientSky, skyMix) * albedo;
outColor = vec4(ambient + direct, 1.0);
}
This is where the last post’s decomposition pays out. The shader has no idea how many triangles were drawn, which mesh they came from, or whether the geometry was static, animated, rasterized or ray traced. It reads the G-Buffer and nothing else, which is what makes the G-Buffer a contract: either side of it can be rewritten without telling the other.
There is one case the geometry pass leaves for this shader to handle, and it is the first consequence of shading in screen space rather than on surfaces. A fullscreen pass runs on every pixel, including the ones no triangle ever touched, and those pixels read a cleared G-Buffer: a zero normal, which is not a direction at all. The shader tests for it and either discards, so the render pass clear color survives, or paints a background of its own. In a forward renderer the case does not exist, because a fragment shader only runs where geometry is. Deferred shading trades that guarantee away and gets a place to draw a sky in return.
The output is HDR, high dynamic range, into an R16G16B16A16Sfloat target.
Color values are allowed past 1.0 because real light intensities go past it.
A bright specular highlight might come out at 5.0, which is physically meaningful and completely undisplayable.
One more transformation is needed before a monitor can show any of it.
Here is what the pass hands on, with the tonemap stepped over and the HDR target put on screen as it stands:

The scene is the one post 3 took apart, now lit by the three buffers that post wrote: same sphere, same cube, same camera. What the picture shows about high dynamic range is where it runs out. A display takes 0 to 1, so every value above 1.0 arrives as the same white as 1.0 itself, and about one lit pixel in six in this frame does exactly that. All of them are on the sphere, and all of them are the red channel: a red surface under a bright key light runs out of range in one channel long before the other two. Read a line straight across the sphere and red is pegged at 255 for its whole width, while green and blue climb from 92 to 163 and fall back. The shape is still in the picture, carried entirely by the two channels that had room left. Nothing in the pass is wrong. The frame is being asked to be a picture before anything has made it into one.
The tonemap pass, and where gamma actually lives
The tonemap pass is the frame’s last transformation: unbounded HDR values in, the [0, 1] range a display can show out. This is the entire shader:
// tonemap.frag, in full.
#version 450
layout(set = 0, binding = 0) uniform sampler2D hdrInput;
layout(location = 0) in vec2 uv;
layout(location = 0) out vec4 outColor;
void main() {
vec3 hdr = texture(hdrInput, uv).rgb;
// Reinhard tonemap. The result is display-referred but still LINEAR light:
// the range is squeezed into 0..1, the transfer function is not applied.
vec3 mapped = hdr / (hdr + vec3(1.0));
outColor = vec4(mapped, 1.0);
}
Reinhard is not the best operator. It desaturates highlights and rolls off in a way that reads flat next to a filmic curve. ACES, Uncharted 2 and AgX are all better for different aesthetic goals, and any of them is a shader edit away. Reinhard is one line, it is honest about what it does, and the pass exists to be swapped.
What is more interesting is the line that is not there.
There is no pow(mapped, 1.0 / 2.2) at the end of this shader, and leaving it out is deliberate.
The pass writes into a target whose format is R8G8B8A8Srgb, and an sRGB format means the hardware applies the encode on write: exactly, piecewise, and after blending rather than before.
Doing it in the shader as well is the classic double encode.
The numbers are worth seeing, because the result does not look like a bug, it looks like a mood.
Mid grey leaves the lighting pass at 0.18, Reinhard brings it to 0.15, the correct encode puts it on screen at 0.42, and a second pass over it lifts that to 0.69.
Every shadow turns milky, the whole frame reads washed out, and nothing anywhere is obviously wrong.
Post 3 arrived at the same rule from the other direction, when the albedo target turned out to want an sRGB format so eight bits would land where the eye needs them.
The transfer function belongs to the format, not to the shader.
Twice now that rule has come up, once on the way into the G-Buffer and once on the way out of the frame, and both times the shader’s job was to do less.
If this pass ever has to write a UNORM target, the encode comes back here, and it comes back as the piecewise curve rather than as pow(1/2.2).
And “swapped later” is exactly why tonemapping is a separate pass at all. You might want a different operator without touching the lighting code, or a bloom pass in between: read the HDR image, extract the bright pixels, blur them, add them back, then tonemap the result. If the tonemap lived at the bottom of the lighting shader, both of those would be edits to lighting. As a pass, each one is additive.
Here is the same frame with the pass back in place:

Nothing in the scene moved, no light changed, and no exposure was touched.
The only difference between this image and the one above is hdr / (hdr + 1), applied at the point where the frame stops being light and starts being a picture.
That same line across the sphere now runs 188 to 204 in red instead of a flat 255, so the channel that had been pegged is carrying the shape again, and no pixel anywhere in the frame is clipped.
The operator’s bill is visible in the same pair.
Every value came down, so the frame is darker than the one above it and the red is less saturated, which is the desaturated roll-off Reinhard is known for, sitting in plain view rather than in a sentence about it.
That is the whole job, and the whole cost: not a prettier image, but an image that fits in the range a display can hold.
The frame is now three passes:
Three functions. Each transforms its inputs into its outputs. None of them knows the others exist.
Which is exactly the problem.
Three passes, three problems
The passes work individually. Composition is what breaks them.
The last post ended on two of these, as the questions the G-Buffer could not answer for itself. Order: lighting runs after the G-Buffer pass and tonemap after lighting, which is obvious at three passes and stops being obvious at ten. Synchronization: GPU execution is asynchronous and pipelined, so without an explicit barrier the lighting shader may sample textures the geometry pass is still writing, and the failure mode is the worst kind, which is that it works on your machine and flickers on someone else’s under load.
The third one post 3 never mentioned, because nothing before this needed it.
Layout transitions.
ColorAttachmentOptimal and ShaderReadOnlyOptimal are not two names for one thing.
They are different physical arrangements of the same pixels, chosen so that writing and sampling are each fast, and the GPU will read one as though it were the other without complaint.
So the layout has to change between the two passes, at the right moment, which means something has to know at every point in the frame which layout every image is in.
You can answer all three by hand: hard-code the order, write the barriers between each pair of passes, track the layouts in your head and in comments.
The engine still contains a pipeline that does exactly this, a standalone G-Buffer pipeline with three hand-written CmdPipelineBarrier calls, and it is correct, and it was not hard to write.
It does not scale.
Add a fourth pass and you are editing barrier code, reordering calls, and re-deriving which layout each image is in at each point.
The work grows as O(passes x resources), and every bit of it is mechanical.
There is no creative decision in transitioning an image from ColorAttachmentWrite to ShaderRead.
It is bookkeeping, and bookkeeping is what machines are for.
What if the passes only declared what they read and write, and something else worked out the rest?
Passes as pure declarations
The render graph starts with a vocabulary for describing passes as data. No execution, no GPU state, no Vulkan types anywhere. Just descriptions.
// RenderLab.Graph/GraphTypes.cs, trimmed of doc comments.
// Logical identity for a resource. Two passes that name the same one are linked.
public readonly record struct ResourceName
{
public string Name { get; }
public static Result<ResourceName, GraphError> Create(string name); // rejects blank
public static ResourceName Of(string name); // throwing, for literals
}
// How a pass touches a resource. This is what drives barrier insertion.
// One more case has been added since: a pass that tests against depth without
// writing it. The last section says which pass wanted it.
public enum ResourceUsage : byte
{
ColorAttachmentWrite,
DepthStencilWrite,
ShaderRead,
Present,
}
public readonly record struct PassInput(ResourceName Resource, ResourceUsage Usage);
public readonly record struct PassOutput(ResourceName Resource, ResourceUsage Usage);
public sealed record RenderPassDeclaration(
string Name,
ImmutableArray<PassInput> Inputs,
ImmutableArray<PassOutput> Outputs);
ResourceName is a value type with a smart constructor rather than a bare string, which is the same move post 1 made for scene values.
A blank resource name is not a resource name, and the type refuses to hold one.
The throwing Of exists for the literals in the pipeline’s own setup, where a failure is a typo in code rather than a fact about input.
With that vocabulary, the whole deferred frame is three records:
// RenderLab.Pipelines/DeferredPipeline.cs, in Initialize, as this post leaves it.
// The frame has since grown a fourth G-Buffer attachment and a fifth pass; the
// last section covers what that did to this listing, which is almost nothing.
var passes = ImmutableArray.Create(
new RenderPassDeclaration("GBuffer",
Inputs: [],
Outputs: [
new PassOutput(gPosition, ResourceUsage.ColorAttachmentWrite),
new PassOutput(gNormal, ResourceUsage.ColorAttachmentWrite),
new PassOutput(gAlbedo, ResourceUsage.ColorAttachmentWrite),
]),
new RenderPassDeclaration("Lighting",
Inputs: [
new PassInput(gPosition, ResourceUsage.ShaderRead),
new PassInput(gNormal, ResourceUsage.ShaderRead),
new PassInput(gAlbedo, ResourceUsage.ShaderRead),
],
Outputs: [new PassOutput(hdrColor, ResourceUsage.ColorAttachmentWrite)]),
new RenderPassDeclaration("Tonemap",
Inputs: [new PassInput(hdrColor, ResourceUsage.ShaderRead)],
Outputs: [new PassOutput(viewport, ResourceUsage.ColorAttachmentWrite)]));
Read those declarations. No pipeline objects, no framebuffers, no command buffers, no shaders. Just: I write these, I read those. Each one is an immutable record you can create, inspect, compare, serialize and test, and it describes intent rather than action.
Look at how the dependencies arrive.
The GBuffer pass writes GBuffer.Position, the Lighting pass reads GBuffer.Position, and that is a dependency that nobody wrote the word “depends” for.
It is implicit in the resource names, exactly the way a functional language derives evaluation order from which expression uses which value.
Which makes a declaration a function signature written as data: inputs are parameters, outputs are return values, the name is the function name, and the graph is a set of signatures that compose because they agree on names. This is what post 1 meant by “rendering is already a data transformation”, made concrete enough to compile.
One detail in that listing is worth stopping on, because it is the strongest evidence I have that the shape is right.
The last pass writes a resource called Viewport, as a color attachment.
It used to write Backbuffer, with usage Present, because the tonemap pass used to draw straight into the swapchain image, which is what a renderer does when the renderer is the whole application.
It is not, any more.
The scene is now a picture that an editor panel samples and draws, and the only pass that touches the swapchain is the interface’s own.
That change altered what the frame’s final image is, which sounds like the sort of thing that rewrites a pipeline.
In the graph it was one line: a different resource name, a different usage.
The compiler re-derived the barrier for it without being asked, because the barrier was never written down in the first place.
The compiler is a pure function
The compiler takes an immutable array of declarations and returns an immutable array of resolved passes: the same declarations, sorted, each carrying the barriers that must be recorded before it runs. Same input, same output, every time. No GPU, no device, no window, no clock.
// RenderLab.Graph/RenderGraphCompiler.cs
public static Result<ImmutableArray<ResolvedPass>, GraphError> Compile(
ImmutableArray<RenderPassDeclaration> passes)
{
return BuildWriterMap(passes)
.Bind(writerByResource => ValidateInputsHaveWriters(passes, writerByResource))
.Bind(writerByResource => TopologicalSort(passes, writerByResource))
.Map(InsertBarriers);
}
Four steps, and the whole function is the pipe between them.
Build the writer map. Walk every pass’s outputs and record which pass writes each resource. If two passes claim the same one, that is not a graph, and the compiler says so rather than picking a winner: a render graph resource has exactly one writer.
Validate the readers. Walk every pass’s inputs and check that something writes what each one reads. A pass reading a resource nobody produces is a typo that otherwise arrives as a black screen and an afternoon.
Sort topologically. This is Kahn’s algorithm (1962), and it is the part people expect a render graph to be made of. Count each pass’s unresolved inputs and start with the passes that have none, which here is the GBuffer pass, since it reads nothing. Emit one, decrement the counts of everything waiting on it, enqueue whatever just reached zero, repeat. If passes are left over when the queue empties, their dependencies are entangled in a cycle and no order satisfies them.
Insert barriers. Walk the sorted passes in order, tracking the last usage of each resource. When a pass touches a resource differently than it was last touched, that difference is a barrier.
// RenderLab.Graph/RenderGraphCompiler.cs, InsertBarriers, trimmed.
var lastUsage = new Dictionary<ResourceName, ResourceUsage>();
foreach (var pass in sortedPasses)
{
var barriers = ImmutableArray.CreateBuilder<BarrierDesc>();
foreach (var input in pass.Inputs)
if (lastUsage.TryGetValue(input.Resource, out var prev) && prev != input.Usage)
barriers.Add(new BarrierDesc(input.Resource, prev, input.Usage));
foreach (var output in pass.Outputs)
if (lastUsage.TryGetValue(output.Resource, out var prev) && prev != output.Usage)
barriers.Add(new BarrierDesc(output.Resource, prev, output.Usage));
resolved.Add(new ResolvedPass(pass, barriers.ToImmutable()));
foreach (var input in pass.Inputs) lastUsage[input.Resource] = input.Usage;
foreach (var output in pass.Outputs) lastUsage[output.Resource] = output.Usage;
}
Outputs get the same treatment as inputs, which is easy to leave out and matters the moment a graph reuses a resource. A pass that reads an image can be followed by a pass that writes it again, and the transition back to a writable layout is as real as the one that made it readable. A blur that ping-pongs between two images does it every frame.
For our three passes the result is: GBuffer with no barriers, Lighting with three (position, normal and albedo, each ColorAttachmentWrite to ShaderRead), and Tonemap with one (HDR, the same transition).
The engine prints the compiled order and the barrier count at startup, and the editor draws the whole thing live, which I will come back to.
The error channel deserves a paragraph of its own, because it is where this compiler differs from the version I first wrote. It does not throw.
public abstract record GraphError
{
public sealed record InvalidResourceName(string Attempted) : GraphError;
public sealed record Cycle(ImmutableArray<string> RemainingPasses) : GraphError;
public sealed record DuplicateWriter(ResourceName Resource, string FirstPass, string SecondPass) : GraphError;
public sealed record UnknownResource(ResourceName Resource, string ConsumerPass) : GraphError;
}
Three ways to describe a frame that cannot exist, each one a value carrying the names involved rather than a message carrying a sentence: Cycle names the passes it could not order, DuplicateWriter names the resource and both claimants, UnknownResource names the resource and the pass that asked for it.
Nothing catches these today, because the deferred pipeline’s declarations are literals, so a failure means the code is wrong and the one call site turns the error into an exception.
What matters is that the compiler’s failure modes are part of its type, not that anything recovers from them today.
A caller that builds a graph from something less trustworthy than a literal, a scene file or a panel where passes get wired by hand, gets to handle them instead of catching them.
Testing a frame without a GPU
This is the payoff I find most satisfying, and the one that is hard to believe until you have it. The frame’s structure is testable. Not the pixels: the structure, which is the part that breaks.
// tests/RenderLab.Graph.Tests/CompilerTests.cs
[Fact]
public void ReversedDeclaration_StillCorrectOrder()
{
var offscreen = ResourceName.Of("OffscreenColor");
var passes = ImmutableArray.Create(
new RenderPassDeclaration("PostProcess",
Inputs: [new PassInput(offscreen, ResourceUsage.ShaderRead)],
Outputs: [new PassOutput(ResourceName.Of("Backbuffer"), ResourceUsage.Present)]),
new RenderPassDeclaration("Geometry",
Inputs: [],
Outputs: [new PassOutput(offscreen, ResourceUsage.ColorAttachmentWrite)]));
var resolved = Ok(RenderGraphCompiler.Compile(passes));
Assert.Equal("Geometry", resolved[0].Declaration.Name);
Assert.Equal("PostProcess", resolved[1].Declaration.Name);
}
The passes are declared backwards on purpose. The compiler sorts them anyway, because declaration order carries no meaning: only the data dependencies do.
Eighteen tests cover the compiler.
A single pass produces no barriers; a two-pass dependency produces the right order and one barrier, checked down to its from and to usages; independent passes both survive.
A diamond, where two passes read one resource and a fourth reads both of theirs, sorts with the fork first and the join last, and the join gets two barriers.
Each of the three error cases comes back as its own record, carrying the pass and resource names that caused it, and three small tests pin the ResourceName constructor.
Every one of them runs in milliseconds, in a plain test project, with no Vulkan, no device, no window, and no GPU in the machine at all. Compare that to testing the equivalent hand-written pipeline: you would need a device, a surface, images, a validation layer, and some way to assert on a stream of recorded commands, and what you would actually do instead is run it and look at the screen.
That is the difference the functional core is for. The frame’s structure became data, and data can be checked.
The executor: one function holds the side effects
The compiler produces a plan. Something has to turn a plan into Vulkan commands, and everything impure about the frame lives in that something.
// RenderLab.Gpu/VulkanGraphExecutor.cs, trimmed.
public static unsafe void Execute(
GpuState state,
CommandBuffer cmd,
ImmutableArray<ResolvedPass> resolvedPasses,
Dictionary<string, Action<Vk, CommandBuffer>> passRecorders,
Dictionary<ResourceName, GraphImage> resourceImages)
{
foreach (var resolved in resolvedPasses)
{
foreach (var barrier in resolved.BarriersBefore)
{
if (!resourceImages.TryGetValue(barrier.Resource, out var target))
continue;
var (srcStage, srcAccess, oldLayout) = MapUsage(barrier.FromUsage);
var (dstStage, dstAccess, newLayout) = MapUsage(barrier.ToUsage);
// ... fill an ImageMemoryBarrier from those six values ...
vk.CmdPipelineBarrier(cmd, srcStage, dstStage, 0, 0, null, 0, null, 1, &imageBarrier);
}
if (passRecorders.TryGetValue(resolved.Declaration.Name, out var recorder))
recorder(vk, cmd);
}
}
Three inputs: the resolved passes, a map from resource names to real Vulkan images, and a map from pass names to recorder functions. Two things happen per pass, in order: the barriers, then the recording.
The barrier translation is a lookup table.
ColorAttachmentWrite means the color attachment output stage, a color write access mask, and ColorAttachmentOptimal.
ShaderRead means the fragment shader stage, a shader read access mask, and ShaderReadOnlyOptimal.
DepthStencilWrite and Present fill in the other two rows.
That table is the entire translation from the pure vocabulary into Vulkan’s, and it is the only place in the engine that knows both.
One wrinkle arrived later and says something about where a boundary belongs.
A barrier covers a subresource range, and for a long time every image the graph moved was a single-mip, single-layer color target, so the executor wrote LevelCount = 1, LayerCount = 1, ColorBit into every barrier it built and nobody noticed the assumption.
Cubemaps broke it: six layers, and a filtered chain of five mips, of which such a barrier transitions exactly one and leaves the other twenty-nine subresources wherever they were.
The fix was a small record on the impure side:
public readonly record struct GraphImage(
Image Image,
uint MipLevels = 1,
uint ArrayLayers = 1,
ImageAspectFlags Aspect = ImageAspectFlags.ColorBit);
The compiler did not change, and could not have. A mip count is a fact about a Vulkan image, not a fact about a dependency between passes. When something new needs describing, the question is which side of the boundary it is a fact about, and that question usually answers itself once it is asked out loud.
Above the executor: declarations, compiler, resolved passes, all pure and all testable. Inside it: barriers, command recording, submission. Draw a line through the architecture and the executor is the line.
It is the same shape post 1 described with a much smaller example, a pure pose function under a thin shell, scaled up to an entire frame. Describe the work as data. Let a pure function organize it. Push the effects to the boundary.
A slot is not a pass
There is a subtlety in Execute that took me a while to see as a feature rather than an accident.
The graph names a pass; it does not contain one.
passRecorders["Tonemap"] is a closure, and what that closure draws is entirely its own business.
The engine uses this immediately. The Visualization panel can put any single G-Buffer target on screen instead of the lit result, and it does that without a fourth pass. The Tonemap slot records a different fragment shader:
// RenderLab.Pipelines/DeferredPipeline.cs, RecordTonemapPass, trimmed.
if (ui.Viz == VisualizationMode.Final)
TonemapPass.Record(api, cb, tonemapResources); // Reinhard over the HDR image
else
DebugVizPass.Record(api, cb, debugVizResources, pc); // one G-Buffer target, raw
The declaration is unchanged either way: read one image, write the viewport. The order is unchanged, the barriers are unchanged, and the picture is completely different. The graph describes the shape of the frame; the recorders decide what gets drawn inside that shape. Those really are two different questions, and the last post’s captures exist because they are.
The compiled graph is also something you can look at. The editor has a panel that walks the resolved passes and prints each one in execution order with its inputs, its outputs and the barriers in front of it, straight out of the same immutable array the executor is iterating. It is read-only, because every value on it was derived and there is nothing there to edit. Making the frame’s structure into data had a consequence nobody planned: the structure became something a tool could show you.
That panel is a list, and a list is the wrong shape for a graph.
It has since become a picture: passes and resources as boxes, what flows between them as lines, the barriers the compiler inserted drawn on the lines that carry them, and each pass sized by what it actually costs on the GPU.
The layout that decides where every box goes is another pure function over the same ResolvedPass array, with its own unit tests and no idea what a pixel is, which means the argument in this section did not change when the drawing arrived.
It just got a better illustration.
I am not going to explain it here, because it is a post of its own and it is the one this whole project is pointed at.
What matters at the end of the introduction is the precondition: you can only draw a pipeline that exists as an artifact, and most renderers do not have one.
The frame, composed
Here is a whole frame, end to end, as the loop actually runs it.
- Compile the graph. Pure, and done once at startup rather than per frame, because the declarations do not change while the app runs.
- Begin the frame. Wait on this frame’s fence, acquire a swapchain image, reset the command buffer. This is the imperative shell at its most imperative.
- Build the interface. The editor’s panels are built first, because building them is what produces the changes the renderer then reads, including how large the viewport is.
- Resize the viewport if it moved. Dragging a panel boundary is a resize of the scene, since the scene renders at the viewport’s size and aspect rather than the window’s.
- Build the scene snapshot.
An immutable
Scenefor this frame: drawables, lights, camera. - Execute the resolved passes. Barriers, then recorders, in the order the compiler decided. Three passes, four barriers, and no hand-written synchronization anywhere in the path that produces the lit image. Both numbers have grown since; neither this step nor the code behind it changed when they did.
- Draw the interface. One overlay pass into the swapchain image, sampling the picture step 6 produced.
- End the frame. Submit, present.
The compiler runs once. The executor runs every frame and does no thinking: it follows a plan computed before the window opened, so the architecture’s runtime cost is a loop over three records and a dictionary lookup per pass.
What we have is a deferred pipeline where the passes do not know about each other, the execution order is derived, the barriers are derived, the layouts are derived, and the entire structure of the frame is testable without a GPU.
What the graph does not do
Every render graph talk you will find describes a system considerably more capable than this one, and it would be dishonest to end without saying where the line is. Each of these is something the graph does not do yet, listed with what it would take.
Depth is not in the graph.
The G-Buffer’s depth attachment is created, used and transitioned entirely outside the compiler’s knowledge, and the deferred pipeline carries exactly one hand-written barrier as a result: the one that makes depth readable when you ask to visualize it.
Depth is the resource whose layouts the render pass object already manages, so adding it means reconciling two owners.
It is the first thing I would fix, and the fix is a DepthStencilWrite declaration plus a GraphImage with a depth aspect.
Since fixed, and worth saying how, because the reason is not the one above. Depth did not enter the graph because I got around to it. It entered when a later pass needed to read it, which is the point at which “who owns this resource” stops being a tidiness question and starts being a barrier you would otherwise write by hand. A resource with one writer and no readers is a resource the compiler can only agree with. The change was the declaration predicted here, plus one case on
ResourceUsagefor testing against depth without writing it.
Resources are not aliased, and lifetimes are not inferred. A real frame graph knows that two temporaries never overlap in time and hands them the same memory; this one allocates every resource for the life of the viewport size. At three passes that costs nothing. At thirty it would be the reason to write the lifetime analysis, and that analysis is a pure function over the same declarations, which is the encouraging part.
Half fixed, in the half I did not expect. The lifetime analysis exists: every resource’s first write and last read, computed from exactly the declarations above, exactly as pure as predicted. What it feeds is not the allocator. It draws a timeline under the graph picture, so you can see that four G-Buffer images are alive at once and three of them are dead before the last pass runs. Aliasing them is still not done, and now the reason is different: the analysis that would justify it is sitting right there, and the frame it would save memory on is not large enough to care. Knowing the shape of a problem is not the same as having the problem.
Nothing is culled. A pass whose outputs nobody reads still runs, where Frostbite’s graph would drop it. Mine would first have to learn that “nobody reads it” and “it is the frame’s final output” are different things, which is a small amount of work on data that is already there.
One queue, no async compute. Everything is recorded into a single command buffer in a single order. Overlapping a compute pass with graphics work means the graph reasoning about queues and semaphores rather than only about barriers, which is a different and much larger piece of work.
The barriers are correct, not minimal.
Each transition is its own vkCmdPipelineBarrier immediately before the pass that needs it, rather than batched with its neighbours or hoisted as early as it could legally go.
Recorders are closures over mutable state. A recorder captures descriptor sets, framebuffers and pipelines that were built imperatively and get rebuilt on resize, which is the shell doing what the shell is for. It does mean “the frame is pure data” is a claim about the frame’s structure, not about every value the frame touches.
The first three are pure functions over declarations I already have; the last three are facts about the boundary and the hardware. The split held up, in the sense that the things that got harder got harder on the side where hard things were supposed to live.
Where the lighting went
The lighting shader in this post is a reduction. The real one has grown a great deal since, and the shape of that growth is the last piece of evidence about the architecture, so it is worth a short tour.
The single light became a buffer of them. Lights are packed into a storage buffer as 48-byte records, a position, a direction, a color and an intensity, with a type tag riding in a spare lane; the pass is told only how many there are, and the shader loops. Point and directional lights differ by that tag and by whether attenuation applies. The ceiling is 64, a number chosen for the buffer rather than by the technique.
The shading model became a runtime choice. Lambert, Phong, Blinn-Phong and Cook-Torrance all live in the same shader behind a mode the editor sets from a dropdown, which turns the comparison between them into an A/B on one frame rather than a rebuild. The ambient term went the same way: a hemispheric constant when there is nothing better, a convolved environment when the scene names one.
The material parameters rode in for free, right up until they did not.
The G-Buffer had two unused alpha channels, one on the normal target and one on albedo, and roughness and metallic went into them: no new attachment, no extra bandwidth, no change to any declaration.
Then the renderer had to read a whole glTF material rather than the half of one it had been reading, and emissive is a colour rather than a scalar.
A colour does not fit in a spare alpha, so the G-Buffer grew a fourth attachment, and that one is a change to the declarations: one more PassOutput on GBuffer, one more PassInput on Lighting.
The push-constant block hit exactly 128 bytes around the same time, which is the minimum every Vulkan implementation guarantees, and the prediction I made about it came true on schedule: the next thing the pass needed to be told arrived in a buffer instead.
One change did more than add a line. Alpha blending cannot go through a G-Buffer at all, because a G-Buffer holds one surface per pixel and a blended surface is by definition not the only one there, so the frame grew a fifth pass that shades transparent geometry the forward way, after the lighting pass has resolved everything opaque. It writes into the HDR image the lighting pass already wrote, which is a third thing a pass can do to a resource: not produce it, not consume it, modify it. The graph had no word for that and now has one. It is also what finally pulled depth into the compiler, for the reason the previous section gave.
Here is what none of it touched: the compiler’s four steps, the executor, the barrier table, or any pass’s implementation other than the one being changed. The declarations moved twice in five milestones, both times because the frame changed shape rather than because a technique did, which is exactly the distinction I wanted the declarations to be sensitive to. Everything else was a fragment shader, a descriptor set, and a buffer upload.
I am not going to promise a post about each of those techniques. Some are worth writing about and some are a paragraph in somebody else’s tutorial. What matters here is the shape of the diff, and the shape of the diff was: almost always the shader got bigger and nothing else moved, and on the two occasions something else did move, the thing that moved was a declaration of what the frame is made of. That is the failure mode I would want if I had to pick one.
Four posts, one frame
This is where the introduction ends, so let me close the loop it opened.
Post 1 was an argument with no code: that rendering is already a data transformation, that engines hide that fact behind objects, and that a functional core with an imperative shell should fit graphics better than graphics usually admits. Post 2 turned on the machine: instance, device, swapchain, render pass, pipeline, command buffer, fence, semaphore, and a triangle as proof that the CPU and the GPU are talking. Post 3 decomposed a frame: forward versus deferred, multiple render targets, and a geometry pass that writes what a surface is and refuses to say anything about light. This post composed it back: a pass that reads that data and lights it, a pass that maps the result onto a display, and a graph that orders and synchronizes them without either one knowing the other exists.
If you have followed all four, you have seen every piece a Vulkan renderer needs between a scene and a pixel. Not every piece a renderer eventually wants. This frame has no shadows, no culling, no anti-aliasing, and a material model that arrived after the fact. But there is no structural hole left in the path: geometry becomes data, data becomes light, light becomes a picture, and a pure function decides the order. Adding a technique to this pipeline is now a local act. Declare a pass, write a shader, register a recorder. The compiler finds the order and the barriers, and the passes that already work do not get edited to make room.
The claim from post 1 held, with a caveat I would rather state than gloss over. The pure core turned out to be small: a handful of record types, a hundred lines of compiler, eighteen tests. The imperative shell turned out to be large, because Vulkan is large, and no amount of architecture makes descriptor pools pleasant. What the split bought was not less Vulkan. It was knowing exactly where the Vulkan is, and being able to answer “is this frame correct?” for the part that decides correctness, in milliseconds, on a machine with no GPU in it.
The frame composes itself. Everything after this is a technique, and a technique is a different kind of post.