git.lucas.co / cce-ui
GPU-accelerated UI toolkit (Vulkan)
git clone https://git.lucas.co/cce-ui.git

commit931caf749aa152fdc17c78013405091b253b12a8
parent0e74e7072b
authorClaude <noreply@anthropic.com>
date2026-10-05 06:49
W4c: the path tracer runs in the browser

Stage3D carries the tracer's half too (set_rt_scene[_with_image],
set_rt_environment, set_rt_background, stage_rt, rt_accumulating), and
what it traces from moves to draw::rt: the schema, the binned-SAH BVH
and its tests, scene packing (pack_scene) and the parameter blocks
(rt_params, denoise_params) as the shaders read them, the constants;
the rt_*.wgsl shaders move to draw/. vk::rt re-exports the schema and
keeps its own state (frames in flight, the ray-query tier,
RtOffscreen), now packing through draw::rt — its GPU tests pass on
lavapipe.

web/rt.rs is the compute tier on WebGPU: the same shaders, packing and
restart rules, a sample a frame and three a-trous iterations in one
compute pass. WebGPU has no ray tracing, so there is no ray-query tier;
and where Vulkan blits the tracer's rgba8unorm image into the sRGB
backdrop, converting as it copies, a small render pass loads each texel
and writes it through the backdrop's sRGB view.

The 3D probe gains a traced mode that stages exactly eight frames:
Vulkan's compute tier on lavapipe against SwiftShader, the traced
pane's mean differs by 0.10 of a level and every pixel is within 8.
Native is unchanged: the same two known test failures (534 pass with
the GPU tests), the plate golden byte-identical, the 24-step run
identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WjL3pejMNY95NHv9BcmXaZ

 CLAUDE.md                         |  27 +-
 Cargo.toml                        |   2 +
 examples/probe3d/scene.rs         |  67 +++-
 examples/probe3d_native.rs        |   8 +-
 examples/probe3d_web.rs           |  16 +-
 scripts/web-probe/capture-app.mjs |  11 +-
 scripts/web-probe/demo.html       |   7 +-
 scripts/web-probe/probe3d         |  17 +-
 src/draw/mod.rs                   |   1 +
 src/draw/rt.rs                    | 691 ++++++++++++++++++++++++++++++++++++++
 src/{vk => draw}/rt_bvh.wgsl      |   0
 src/{vk => draw}/rt_common.wgsl   |   0
 src/{vk => draw}/rt_denoise.wgsl  |   0
 src/{vk => draw}/rt_query.wgsl    |   0
 src/draw/scene.rs                 |  28 +-
 src/draw/shaders.rs               |  13 +
 src/engine.rs                     |   1 +
 src/vk/renderer.rs                |  15 +
 src/vk/rt.rs                      | 686 ++-----------------------------------
 src/web/mod.rs                    |   1 +
 src/web/renderer.rs               |  56 ++-
 src/web/rt.rs                     | 469 ++++++++++++++++++++++++++
 22 files changed, 1437 insertions(+), 679 deletions(-)

diff --git a/CLAUDE.md b/CLAUDE.md
index ff7c48c..81a71bc 100644
--- a/CLAUDE.md
+++ b/CLAUDE.md
@@ -162,6 +162,28 @@ image in the scene before the translucent draw, a host light, frost over the pan
 195 px differ by more than 8 levels, all on 1 px wires (where along its length a line
 steps a row is the rasterizer's), everything else within 2.
 
+**And so does the path tracer** (since 2026-10-05). `Stage3D` carries the tracer's half
+too — `set_rt_scene` / `set_rt_scene_with_image`, `set_rt_environment`,
+`set_rt_background`, `stage_rt`, `rt_accumulating` — and what it traces from moved to
+`draw::rt`: the schema (`RtTriangle`, `RtMaterial`, `RtImage`, `RtCamera`,
+`RtEnvironment`), the binned-SAH BVH and its tests, the buffers (`pack_scene`) and
+parameter blocks (`rt_params`, `denoise_params`) as the shaders read them, and the
+constants; the shaders (`rt_common` / `rt_bvh` / `rt_query` / `rt_denoise`) moved to `draw/`.
+`vk::rt` re-exports the schema and keeps its own state (frames in flight, the ray-query
+tier, `RtOffscreen`) — it now packs and lays out through `draw::rt`, verified by its
+GPU tests (`cargo test --lib rt -- --ignored`, lavapipe). `web/rt.rs` is the compute tier on
+WebGPU: the same shaders and packing, the same rules for restarting the accumulation, one
+sample a frame plus the three à-trous iterations in one compute pass. Two differences:
+WebGPU has no ray tracing, so there is no ray-query tier (the compute tier is what every
+Vulkan device without RT cores runs too); and Vulkan BLITS the tracer's `rgba8unorm`
+image into the sRGB backdrop, converting as it copies, which WebGPU's copies cannot — a
+small render pass loads each texel and writes it through the backdrop's sRGB view, the
+same conversion. The probe's traced mode (`Probe3d<true>`, `PROBE3D_TRACE=1` natively,
+`scripts/web-probe/probe3d <out> traced`) stages exactly eight frames and stops, so both
+are compared at eight samples: 2026-10-05, Vulkan compute tier (`CCE_VK_RT=compute`) on
+lavapipe vs SwiftShader, the traced pane's mean differs by 0.10 of a level, every pixel
+within 8, 47 channels in the frame past 8 — a few paths that diverged.
+
 **The reference app runs on both, through one input script.** `examples/demo_web.rs` is
 `src/main.rs`'s `DemoApp` (included by `#[path]`, hence `pub(crate)`) in a page;
 `scripts/web-probe/demo <dir>` builds it, serves it with the machine's fonts and replays the
@@ -1270,7 +1292,9 @@ cce-system-interface) to confirm behavior, not just the test suite.
   at a per-batch dynamic offset (`WEBGPU_BLOCK_STRIDE`); the backdrop is sampled with
   `textureSampleLevel(…, 0.0)` because WebGPU rejects implicit-LOD sampling in the
   non-uniform blur branch (the backdrop has one level, so the texel is the same —
-  `frost_pair` is identical to the pixel either way).
+  `frost_pair` is identical to the pixel either way). And the 3D halves: `draw::scene`
+  (the raster scene's types, uniforms and the `Stage3D` trait) and `draw::rt` (the path
+  tracer's schema, BVH and parameter blocks), with their shaders beside the 2D ones.
 - `web/` — wasm32 only: `WebRenderer` (`new(canvas).await`, `resize`, `prepare_text`,
   `draw_frame_2d`, and `capture_next_frame` / `take_capture().await` or
   `take_pending_capture` to read a frame back). Its module doc lists what differs from the
@@ -1278,6 +1302,7 @@ cce-system-interface) to confirm behavior, not just the test suite.
   dynamic-offset uniform, a 1x1 backdrop, the blur snapshot as end-pass / copy / resume,
   every frame drawn whole. And `shell.rs`, the browser shell: `run`, `Fonts`, `Sizing`,
   `capture` (see "And an `Application` runs in a page" above); `scene.rs`, the 3D pass;
+  `rt.rs`, the path tracer's compute tier;
   `compute.rs`, the async
   `ComputeDevice`; and `request_device`, the adapter and device every one of them asks
   for (with the limits a caller names raised to the adapter's).
diff --git a/Cargo.toml b/Cargo.toml
index 3086395..962a607 100644
--- a/Cargo.toml
+++ b/Cargo.toml
@@ -179,6 +179,8 @@ web-sys = { version = "0.3", features = [
     "GpuCullMode",
     "GpuFrontFace",
     "GpuRenderPassDepthStencilAttachment",
+    "GpuStorageTextureAccess",
+    "GpuStorageTextureBindingLayout",
 ] }
 
 # The renderer probe: one scene through both renderers (examples/probe/).
diff --git a/examples/probe3d/scene.rs b/examples/probe3d/scene.rs
index 028c116..0f460d8 100644
--- a/examples/probe3d/scene.rs
+++ b/examples/probe3d/scene.rs
@@ -12,7 +12,8 @@
 //! blur samples the backdrop the scene left.
 
 use cce_ui::engine::{
-    AppSender, Application, LogicalPosition, LogicalSize, MeshId, SceneDraw, SceneImage, Stage3D, Vertex3D, WindowSettings,
+    AppSender, Application, LogicalPosition, LogicalSize, MeshId, RtCamera, RtImage, RtMaterial, RtTriangle, SceneDraw,
+    SceneImage, Stage3D, Vertex3D, WindowSettings,
 };
 use cce_ui::scene::layout::Rect;
 use cce_ui::scene::paint::{DisplayList, PaintCtx, PlateSpec};
@@ -25,11 +26,22 @@ pub const H: u32 = 800;
 /// The UI strip's width; the 3D pane is the rest of the window.
 const STRIP: f32 = 300.0;
 
-pub struct Probe3d {
+/// The traced probe's sample count: frames staged before it stops.
+pub const TRACE_FRAMES: u32 = 8;
+
+pub struct Probe3d<const TRACE: bool> {
     image: u32,
     meshes: Option<Meshes>,
+    /// Traced frames staged so far.
+    traced: u32,
 }
 
+pub type Raster = Probe3d<false>;
+pub type Traced = Probe3d<true>;
+
+/// The image's corners in the scene, as both passes place it.
+const IMAGE_CORNERS: [[f32; 3]; 4] = [[-2.6, 1.6, -1.6], [-0.6, 1.6, -1.6], [-0.6, 0.35, -1.6], [-2.6, 0.35, -1.6]];
+
 struct Meshes {
     background: MeshId,
     cube: MeshId,
@@ -112,7 +124,25 @@ fn sphere(center: Vec3, radius: f32, stacks: u32, slices: u32) -> (Vec<Vertex3D>
     (tris, lines)
 }
 
-impl Probe3d {
+/// The traced scene: the raster meshes' triangles, a material each.
+fn traced_scene() -> (Vec<RtTriangle>, Vec<RtMaterial>) {
+    let mut tris = Vec::new();
+    let mut mats = Vec::new();
+    let mut add = |verts: Vec<Vertex3D>, albedo: [f32; 3], emission: [f32; 3]| {
+        let m = mats.len() as u32;
+        mats.push(RtMaterial { albedo, emission });
+        for t in verts.chunks_exact(3) {
+            tris.push(RtTriangle { p0: t[0].position, p1: t[1].position, p2: t[2].position, material: m });
+        }
+    };
+    add(cuboid(Vec3::new(0.0, -0.75, 0.0), Vec3::new(3.0, 0.05, 2.4), [[0.3; 3]; 6]), [0.55, 0.57, 0.55], [0.0; 3]);
+    add(cuboid(Vec3::new(-1.6, 0.0, 0.2), Vec3::splat(0.6), [[0.0; 3]; 6]), [0.8, 0.35, 0.25], [0.0; 3]);
+    add(sphere(Vec3::new(0.4, 0.3, -0.4), 0.9, 12, 20).0, [0.45, 0.55, 0.75], [0.0; 3]);
+    add(cuboid(Vec3::new(1.5, 0.1, 1.2), Vec3::splat(0.55), [[0.0; 3]; 6]), [0.3, 0.7, 0.9], [0.4, 0.9, 1.1]);
+    (tris, mats)
+}
+
+impl<const TRACE: bool> Probe3d<TRACE> {
     fn camera(size: LogicalSize, scale: f64) -> (Mat4, (u32, u32, u32, u32)) {
         let s = scale as f32;
         let (pw, ph) = ((size.width - STRIP) * s, size.height * s);
@@ -124,9 +154,19 @@ impl Probe3d {
         let shift = Mat4::from_translation(Vec3::new(1.0 - ndc_w, 0.0, 0.0)) * Mat4::from_scale(Vec3::new(ndc_w, 1.0, 1.0));
         (shift * proj * view, ((STRIP * s) as u32, 0, pw as u32, ph as u32))
     }
+
+    /// The traced pane's camera: the pane's own projection, unshifted —
+    /// the tracer's image is the pane.
+    fn trace_camera(size: LogicalSize, scale: f64) -> RtCamera {
+        let s = scale as f32;
+        let (pw, ph) = ((size.width - STRIP) * s, size.height * s);
+        let proj = Mat4::perspective_rh(40f32.to_radians(), pw / ph, 0.1, 100.0);
+        let view = Mat4::look_at_rh(Vec3::new(3.2, 2.4, 5.6), Vec3::new(0.0, 0.2, 0.0), Vec3::Y);
+        RtCamera { inv_mvp: (proj * view).inverse().to_cols_array_2d() }
+    }
 }
 
-impl Application for Probe3d {
+impl<const TRACE: bool> Application for Probe3d<TRACE> {
     type Message = ();
 
     fn create(_sender: AppSender<()>) -> Self {
@@ -138,7 +178,7 @@ impl Application for Probe3d {
                 px.extend([(x * 255 / (iw - 1)) as u8, (y * 255 / (ih - 1)) as u8, 200, a]);
             }
         }
-        Probe3d { image: cce_ui::draw::upload_rgba(px, iw, ih), meshes: None }
+        Probe3d { image: cce_ui::draw::upload_rgba(px, iw, ih), meshes: None, traced: 0 }
     }
 
     fn settings(&self) -> WindowSettings {
@@ -181,9 +221,24 @@ impl Application for Probe3d {
         let glass_wires = stage.create_mesh(&cuboid_edges(glass_c, Vec3::splat(0.55), [0.9, 0.95, 1.0]));
         stage.set_scene_light([0.6, 0.7, 0.4]);
         self.meshes = Some(Meshes { background, cube, prelit, sphere, sphere_wires, glass, glass_wires });
+        if TRACE {
+            let (tris, mats) = traced_scene();
+            stage.set_rt_scene_with_image(&tris, &mats, Some(RtImage { image: self.image, corners: IMAGE_CORNERS, opacity: 0.9 }));
+        }
     }
 
     fn stage_3d(&mut self, stage: &mut dyn Stage3D, size: LogicalSize, scale: f64) -> bool {
+        if TRACE {
+            // A sample a frame for TRACE_FRAMES frames, then the backdrop
+            // keeps the result.
+            if self.traced >= TRACE_FRAMES {
+                return false;
+            }
+            let (_, pane) = Self::camera(size, scale);
+            stage.stage_rt(pane, Self::trace_camera(size, scale));
+            self.traced += 1;
+            return self.traced < TRACE_FRAMES;
+        }
         let Some(m) = &self.meshes else { return false };
         let (mvp, scissor) = Self::camera(size, scale);
         let mvp = mvp.to_cols_array_2d();
@@ -210,7 +265,7 @@ impl Application for Probe3d {
         stage.stage_scene(scissor, draws);
         stage.stage_scene_images(vec![SceneImage {
             image: self.image,
-            corners: [[-2.6, 1.6, -1.6], [-0.6, 1.6, -1.6], [-0.6, 0.35, -1.6], [-2.6, 0.35, -1.6]],
+            corners: IMAGE_CORNERS,
             mvp,
             opacity: 0.9,
             before: 5,
diff --git a/examples/probe3d_native.rs b/examples/probe3d_native.rs
index 5d3935f..fd91f84 100644
--- a/examples/probe3d_native.rs
+++ b/examples/probe3d_native.rs
@@ -1,9 +1,15 @@
 //! The 3D probe (`probe3d/scene.rs`) on the Vulkan renderer, in the Wayland
 //! runner: screenshot it and compare it with `probe3d_web`'s frame.
+//! `PROBE3D_TRACE=1` runs the traced probe (`CCE_VK_RT=compute` keeps
+//! Vulkan on the compute tier the browser has).
 
 #[path = "probe3d/scene.rs"]
 mod scene;
 
 fn main() {
-    cce_ui::engine::run::<scene::Probe3d>();
+    if std::env::var("PROBE3D_TRACE").is_ok_and(|v| v == "1") {
+        cce_ui::engine::run::<scene::Traced>();
+    } else {
+        cce_ui::engine::run::<scene::Raster>();
+    }
 }
diff --git a/examples/probe3d_web.rs b/examples/probe3d_web.rs
index 139e41e..a868f52 100644
--- a/examples/probe3d_web.rs
+++ b/examples/probe3d_web.rs
@@ -12,11 +12,21 @@ mod web {
     use cce_ui::web::{Fonts, Sizing};
     use wasm_bindgen::prelude::*;
 
+    fn fonts(fonts: js_sys::Array) -> Fonts {
+        std::panic::set_hook(Box::new(|info| web_sys::console::error_1(&info.to_string().into())));
+        Fonts::new(fonts.iter().map(|f| js_sys::Uint8Array::new(&f).to_vec()).collect())
+    }
+
+    /// The raster probe.
     #[wasm_bindgen]
     pub async fn start(canvas: web_sys::HtmlCanvasElement, fonts: js_sys::Array, _families: String, _fill: bool) -> Result<(), JsValue> {
-        std::panic::set_hook(Box::new(|info| web_sys::console::error_1(&info.to_string().into())));
-        let fonts = Fonts::new(fonts.iter().map(|f| js_sys::Uint8Array::new(&f).to_vec()).collect());
-        cce_ui::web::run::<super::scene::Probe3d>(canvas, fonts, Sizing::Page).await
+        cce_ui::web::run::<super::scene::Raster>(canvas, self::fonts(fonts), Sizing::Page).await
+    }
+
+    /// The traced probe (`demo.html?app=probe3d_web&entry=start_traced`).
+    #[wasm_bindgen]
+    pub async fn start_traced(canvas: web_sys::HtmlCanvasElement, fonts: js_sys::Array, _families: String, _fill: bool) -> Result<(), JsValue> {
+        cce_ui::web::run::<super::scene::Traced>(canvas, self::fonts(fonts), Sizing::Page).await
     }
 
     #[wasm_bindgen]
diff --git a/scripts/web-probe/capture-app.mjs b/scripts/web-probe/capture-app.mjs
index 7b62ee6..ecff41a 100644
--- a/scripts/web-probe/capture-app.mjs
+++ b/scripts/web-probe/capture-app.mjs
@@ -1,13 +1,14 @@
-// capture-app.mjs <site-dir> <app> <out.rgba> [settle-ms]: open demo.html
-// with ?app=<app> (browser.mjs), let the app run for settle-ms (default
-// 3000), and write the frame `capture()` reads back from the GPU to
+// capture-app.mjs <site-dir> <app> <out.rgba> [settle-ms] [entry]: open
+// demo.html with ?app=<app> (and &entry=<entry>, the app's start function in
+// place of `start`) through browser.mjs, let the app run for settle-ms
+// (default 3000), and write the frame `capture()` reads back from the GPU to
 // <out.rgba> (premultiplied RGBA8, 1280x800).
 import fs from 'fs';
 import { open } from './browser.mjs';
 
-const [root, app, out, settle] = process.argv.slice(2);
+const [root, app, out, settle, entry] = process.argv.slice(2);
 if (!root || !app || !out) { console.error('usage: capture-app.mjs <site-dir> <app> <out.rgba> [settle-ms]'); process.exit(2); }
-const { page, logs, close } = await open(root, 'demo.html', 1280, 800, `?app=${app}`);
+const { page, logs, close } = await open(root, 'demo.html', 1280, 800, `?app=${app}` + (entry ? `&entry=${entry}` : ''));
 let ok = true;
 try {
   await page.waitForFunction(() => window.demoReady === true, null, { timeout: 120000 });
diff --git a/scripts/web-probe/demo.html b/scripts/web-probe/demo.html
index a01c1b9..f1813c7 100644
--- a/scripts/web-probe/demo.html
+++ b/scripts/web-probe/demo.html
@@ -3,7 +3,8 @@
 <title>cce-ui</title>
 <!-- An app on the browser shell: examples/demo_web.rs, or the example the
      query names (?app=probe3d_web) — any that exports start(canvas, fonts,
-     families, fill) and capture(). The page owns the canvas's size (fill =
+     families, fill), or the entry the query names in its place
+     (&entry=start_traced), and capture(). The page owns the canvas's size (fill =
      Sizing::Page). window.capture() reads the next frame back from the GPU,
      base64 RGBA8: a headless Chromium leaves a WebGPU canvas out of its
      screenshots. -->
@@ -11,7 +12,9 @@
 <canvas id="c"></canvas>
 <script type="module">
 const app = new URLSearchParams(location.search).get('app') || 'demo_web';
-const { default: init, start, capture } = await import(`./${app}.js`);
+const mod = await import(`./${app}.js`);
+const { default: init, capture } = mod;
+const start = mod[new URLSearchParams(location.search).get('entry') || 'start'];
 await init();
 const { files, families } = await (await fetch('fonts.json')).json();
 const fonts = [];
diff --git a/scripts/web-probe/probe3d b/scripts/web-probe/probe3d
index e117efe..a1494ef 100755
--- a/scripts/web-probe/probe3d
+++ b/scripts/web-probe/probe3d
@@ -1,7 +1,9 @@
 #!/usr/bin/env bash
-# web-probe/probe3d <out.rgba> — run the 3D probe (examples/probe3d/scene.rs)
-# on the browser shell and WebGPU renderer in headless Chromium, and write a
-# frame read back from the GPU: 1280x800 premultiplied RGBA8.
+# web-probe/probe3d <out.rgba> [traced] — run the 3D probe
+# (examples/probe3d/scene.rs) on the browser shell and WebGPU renderer in
+# headless Chromium, and write a frame read back from the GPU: 1280x800
+# premultiplied RGBA8. With `traced`, the path-traced probe, given time to
+# take its eight samples ($PROBE3D_SETTLE ms, default 60000).
 #
 # Its other half is `cargo run --example probe3d_native`, screenshotted in a
 # compositor; compare.py composites this frame over that compositor's
@@ -10,7 +12,8 @@
 # `run` needs.
 set -euo pipefail
 here=$(cd "$(dirname "$0")" && pwd)
-out=$(realpath -m "${1:?usage: web-probe/probe3d <out.rgba>}")
+out=$(realpath -m "${1:?usage: web-probe/probe3d <out.rgba> [traced]}")
+mode=${2:-}
 fonts=${PROBE_FONTS_DIR:-/usr/share/fonts/truetype/dejavu}
 cd "$here/../.."
 cargo build --release --target wasm32-unknown-unknown --example probe3d_web
@@ -22,5 +25,9 @@ cp "$here/demo.html" "$site/demo.html"
 mkdir "$site/fonts"
 cp "$fonts"/*.ttf "$site/fonts/"
 (cd "$site/fonts" && printf '%s\n' *.ttf | python3 -c 'import json,sys; print(json.dumps({"files": sys.stdin.read().split(), "families": ""}))') > "$site/fonts.json"
-node "$here/capture-app.mjs" "$site" probe3d_web "$out"
+if [ "$mode" = traced ]; then
+    node "$here/capture-app.mjs" "$site" probe3d_web "$out" "${PROBE3D_SETTLE:-60000}" start_traced
+else
+    node "$here/capture-app.mjs" "$site" probe3d_web "$out"
+fi
 echo "web-probe: wrote $out"
diff --git a/src/draw/mod.rs b/src/draw/mod.rs
index 50270ce..b096ce1 100644
--- a/src/draw/mod.rs
+++ b/src/draw/mod.rs
@@ -14,6 +14,7 @@ use cosmic_text::Buffer as TextBuffer;
 
 pub mod glyphs;
 pub mod images;
+pub mod rt;
 pub mod scene;
 pub mod shaders;
 
diff --git a/src/draw/rt.rs b/src/draw/rt.rs
new file mode 100644
index 0000000..4845067
--- /dev/null
+++ b/src/draw/rt.rs
@@ -0,0 +1,691 @@
+//! The path tracer's scene and the work done on the CPU for it, with no GPU
+//! in it: the schema an app fills ([`RtTriangle`], [`RtMaterial`],
+//! [`RtImage`], [`RtCamera`], [`RtEnvironment`]), the binned-SAH BVH built
+//! over it, the buffers and parameter blocks laid out as `rt_common.wgsl` /
+//! `rt_bvh.wgsl` / `rt_denoise.wgsl` read them, and the tracer's constants.
+//! The Vulkan stage (`vk::rt`, with a hardware ray-query tier besides) and
+//! the WebGPU one (`web::rt`, the compute tier) both trace from these.
+//!
+//! The tracer is plain compute: a CPU-built BVH traversed per pixel,
+//! progressive accumulation of one sample a frame, an à-trous denoise over
+//! the running mean, blitted into the backdrop's pane region in place of the
+//! raster scene. Scene schema is internal by design — importers (OBJ/glTF)
+//! belong in a loader that converts *into* [`RtTriangle`]/[`RtMaterial`].
+
+/// One triangle of an RT scene, in the same space as the camera's `inv_mvp`
+/// (for the designer: mesh space, the space `Vertex3D` positions live in).
+#[derive(Debug, Clone, Copy, PartialEq)]
+pub struct RtTriangle {
+    pub p0: [f32; 3],
+    pub p1: [f32; 3],
+    pub p2: [f32; 3],
+    /// Index into the material slice passed alongside.
+    pub material: u32,
+}
+
+/// Lambertian surface + optional emission, linear color (matching the raster
+/// path, whose vertex colors land in the sRGB attachment as linear values).
+#[derive(Debug, Clone, Copy, PartialEq)]
+pub struct RtMaterial {
+    pub albedo: [f32; 3],
+    pub emission: [f32; 3],
+}
+
+/// An image standing in the traced scene: a quad that shows the image's
+/// colour as it is — unlit, as the raster pass's `SceneImage` draws it — and
+/// lets a ray through where the image is clear. A path ends on it, so to
+/// the rest of the scene it is a light of its own colour. The tracer adds
+/// the quad to the scene itself.
+///
+/// The image is one the 2D pass already holds (an id from `upload_rgba`),
+/// so a picture shown in the raster viewport costs the tracer nothing more.
+#[derive(Debug, Clone, Copy, PartialEq)]
+pub struct RtImage {
+    pub image: u32,
+    /// The quad's corners in the scene's space, in the image's own order:
+    /// top-left, top-right, bottom-right, bottom-left.
+    pub corners: [[f32; 3]; 4],
+    /// Alpha multiplier over the image's own.
+    pub opacity: f32,
+}
+
+/// [`RtImage`] for the headless tracer, which has no 2D pass to share an
+/// image with and is handed the pixels: tightly packed sRGB RGBA8.
+#[derive(Debug, Clone, Copy)]
+pub struct RtImagePixels<'a> {
+    pub pixels: &'a [u8],
+    pub width: u32,
+    pub height: u32,
+    pub corners: [[f32; 3]; 4],
+    pub opacity: f32,
+}
+
+/// The full camera: the inverse of the raster path's `proj * view * model`.
+/// Rays are unprojected from NDC through it, so any matrix stack that renders
+/// the raster viewport drives the tracer unchanged.
+#[derive(Debug, Clone, Copy, PartialEq)]
+pub struct RtCamera {
+    pub inv_mvp: [[f32; 4]; 4],
+}
+
+// --- GPU layouts (must match rt.wgsl) ---
+
+#[repr(C)]
+#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
+pub(crate) struct GpuTriangle {
+    pub(crate) p0: [f32; 4], // w = material index (bitcast)
+    pub(crate) p1: [f32; 4],
+    pub(crate) p2: [f32; 4],
+}
+
+#[repr(C)]
+#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
+pub(crate) struct GpuMaterial {
+    pub(crate) albedo: [f32; 4],
+    pub(crate) emission: [f32; 4],
+}
+
+#[repr(C)]
+#[derive(Debug, Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
+pub(crate) struct GpuBvhNode {
+    pub(crate) min: [f32; 3],
+    /// Leaf (`count > 0`): first triangle. Internal: left child; right = +1.
+    pub(crate) left_first: u32,
+    pub(crate) max: [f32; 3],
+    pub(crate) count: u32,
+}
+
+#[repr(C)]
+#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
+pub(crate) struct RtParams {
+    pub(crate) inv_mvp: [[f32; 4]; 4],
+    pub(crate) width: u32,
+    pub(crate) height: u32,
+    pub(crate) sample_index: u32,
+    pub(crate) max_bounces: u32,
+    pub(crate) spp: u32,
+    pub(crate) _pad: [u32; 3],
+    pub(crate) img_origin: [f32; 4],
+    pub(crate) img_u: [f32; 4],
+    pub(crate) img_v: [f32; 4],
+    pub(crate) background: [f32; 4],
+    // The environment, each xyz in a vec4 (w unused): toward the sun
+    // (unit), the sun's radiance, the sky overhead, the sky below.
+    pub(crate) sun_dir: [f32; 4],
+    pub(crate) sun_color: [f32; 4],
+    pub(crate) sky_zenith: [f32; 4],
+    pub(crate) sky_nadir: [f32; 4],
+}
+
+/// The traced scene's light: a sky that grades from `sky_nadir` straight
+/// down to `sky_zenith` straight up, and a sun — a bright lobe toward
+/// `sun_direction` of `sun_color` radiance. It is the tracer's ONLY light
+/// (nothing in a scene emits unless its material does). Colours are linear
+/// RGB and may exceed 1. [`Default`] is the studio sky the tracer always
+/// had.
+#[derive(Clone, Copy, Debug, PartialEq)]
+pub struct RtEnvironment {
+    /// Toward the sun, world space; any length (zero is straight up).
+    pub sun_direction: [f32; 3],
+    pub sun_color: [f32; 3],
+    pub sky_zenith: [f32; 3],
+    pub sky_nadir: [f32; 3],
+}
+
+impl Default for RtEnvironment {
+    fn default() -> Self {
+        Self {
+            sun_direction: [0.45, 0.75, 0.35],
+            sun_color: [8.0, 7.6, 6.8],
+            sky_zenith: [0.72, 0.82, 0.98],
+            sky_nadir: [0.32, 0.31, 0.35],
+        }
+    }
+}
+
+impl RtEnvironment {
+    pub(crate) fn sun_unit(&self) -> [f32; 3] {
+        let v = glam::Vec3::from_array(self.sun_direction);
+        if v.length_squared() > 1e-12 { v.normalize().to_array() } else { [0.0, 1.0, 0.0] }
+    }
+}
+
+// --- BVH construction (binned SAH) ---
+
+const BVH_BINS: usize = 8;
+const BVH_LEAF_MAX: u32 = 4;
+
+#[derive(Clone, Copy)]
+struct Aabb {
+    min: [f32; 3],
+    max: [f32; 3],
+}
+
+impl Aabb {
+    const EMPTY: Aabb = Aabb { min: [f32::INFINITY; 3], max: [f32::NEG_INFINITY; 3] };
+
+    fn grow(&mut self, p: [f32; 3]) {
+        for a in 0..3 {
+            self.min[a] = self.min[a].min(p[a]);
+            self.max[a] = self.max[a].max(p[a]);
+        }
+    }
+
+    fn grow_aabb(&mut self, other: &Aabb) {
+        self.grow(other.min);
+        self.grow(other.max);
+    }
+
+    fn half_area(&self) -> f32 {
+        let dx = (self.max[0] - self.min[0]).max(0.0);
+        let dy = (self.max[1] - self.min[1]).max(0.0);
+        let dz = (self.max[2] - self.min[2]).max(0.0);
+        dx * dy + dy * dz + dz * dx
+    }
+}
+
+fn tri_aabb(t: &RtTriangle) -> Aabb {
+    let mut b = Aabb::EMPTY;
+    b.grow(t.p0);
+    b.grow(t.p1);
+    b.grow(t.p2);
+    b
+}
+
+fn tri_centroid(t: &RtTriangle) -> [f32; 3] {
+    let mut c = [0.0f32; 3];
+    for a in 0..3 {
+        c[a] = (t.p0[a] + t.p1[a] + t.p2[a]) / 3.0;
+    }
+    c
+}
+
+/// Build a BVH over `triangles`, reordering them so leaves reference
+/// contiguous ranges. Returns the flat node array (empty input → empty vec).
+pub(crate) fn build_bvh(triangles: &mut Vec<RtTriangle>) -> Vec<GpuBvhNode> {
+    if triangles.is_empty() {
+        return Vec::new();
+    }
+    let bounds: Vec<Aabb> = triangles.iter().map(tri_aabb).collect();
+    let centroids: Vec<[f32; 3]> = triangles.iter().map(tri_centroid).collect();
+    let mut order: Vec<u32> = (0..triangles.len() as u32).collect();
+
+    fn range_bounds(order: &[u32], bounds: &[Aabb]) -> Aabb {
+        let mut b = Aabb::EMPTY;
+        for &i in order {
+            b.grow_aabb(&bounds[i as usize]);
+        }
+        b
+    }
+
+    let mut nodes: Vec<GpuBvhNode> = Vec::with_capacity(triangles.len() * 2);
+    let root_bounds = range_bounds(&order, &bounds);
+    nodes.push(GpuBvhNode {
+        min: root_bounds.min,
+        left_first: 0,
+        max: root_bounds.max,
+        count: triangles.len() as u32,
+    });
+
+    // (node index, start, count) work list over `order`.
+    let mut work = vec![(0usize, 0usize, triangles.len())];
+    while let Some((node_idx, start, count)) = work.pop() {
+        if (count as u32) <= BVH_LEAF_MAX {
+            continue; // stays a leaf
+        }
+        let slice = &mut order[start..start + count];
+
+        // Centroid bounds pick the split axis.
+        let mut cb = Aabb::EMPTY;
+        for &i in slice.iter() {
+            cb.grow(centroids[i as usize]);
+        }
+        let mut axis = 0;
+        let mut extent = 0.0f32;
+        for a in 0..3 {
+            let e = cb.max[a] - cb.min[a];
+            if e > extent {
+                extent = e;
+                axis = a;
+            }
+        }
+
+        let mut split_at = None;
+        if extent > 1e-12 {
+            // Binned SAH along `axis`.
+            let scale = BVH_BINS as f32 / extent;
+            let bin_of = |i: u32| -> usize {
+                (((centroids[i as usize][axis] - cb.min[axis]) * scale) as usize)
+                    .min(BVH_BINS - 1)
+            };
+            let mut bin_bounds = [Aabb::EMPTY; BVH_BINS];
+            let mut bin_counts = [0usize; BVH_BINS];
+            for &i in slice.iter() {
+                let b = bin_of(i);
+                bin_counts[b] += 1;
+                bin_bounds[b].grow_aabb(&bounds[i as usize]);
+            }
+            // Cost of each of the BINS-1 split planes.
+            let mut best_cost = f32::INFINITY;
+            let mut best_plane = 0usize;
+            for plane in 1..BVH_BINS {
+                let (mut lb, mut rb) = (Aabb::EMPTY, Aabb::EMPTY);
+                let (mut lc, mut rc) = (0usize, 0usize);
+                for b in 0..plane {
+                    lb.grow_aabb(&bin_bounds[b]);
+                    lc += bin_counts[b];
+                }
+                for b in plane..BVH_BINS {
+                    rb.grow_aabb(&bin_bounds[b]);
+                    rc += bin_counts[b];
+                }
+                if lc == 0 || rc == 0 {
+                    continue;
+                }
+                let cost = lb.half_area() * lc as f32 + rb.half_area() * rc as f32;
+                if cost < best_cost {
+                    best_cost = cost;
+                    best_plane = plane;
+                }
+            }
+            if best_plane > 0 {
+                let mut mid = 0usize;
+                for k in 0..count {
+                    if bin_of(slice[k]) < best_plane {
+                        slice.swap(k, mid);
+                        mid += 1;
+                    }
+                }
+                if mid > 0 && mid < count {
+                    split_at = Some(mid);
+                }
+            }
+        }
+        // Degenerate centroids or a one-sided SAH result: median split keeps
+        // the tree balanced instead of forcing a giant leaf.
+        let mid = split_at.unwrap_or(count / 2);
+
+        let left_bounds = range_bounds(&slice[..mid], &bounds);
+        let right_bounds = range_bounds(&slice[mid..], &bounds);
+        let left_idx = nodes.len();
+        nodes.push(GpuBvhNode {
+            min: left_bounds.min,
+            left_first: (start) as u32,
+            max: left_bounds.max,
+            count: mid as u32,
+        });
+        nodes.push(GpuBvhNode {
+            min: right_bounds.min,
+            left_first: (start + mid) as u32,
+            max: right_bounds.max,
+            count: (count - mid) as u32,
+        });
+        nodes[node_idx].left_first = left_idx as u32;
+        nodes[node_idx].count = 0;
+        work.push((left_idx, start, mid));
+        work.push((left_idx + 1, start + mid, count - mid));
+    }
+
+    // Apply the final order to the triangle array so leaf ranges are direct.
+    let reordered: Vec<RtTriangle> =
+        order.iter().map(|&i| triangles[i as usize]).collect();
+    *triangles = reordered;
+    nodes
+}
+
+// --- The Vulkan stage ---
+
+pub(crate) const MAX_SAMPLES: u32 = 1024;
+pub(crate) const MAX_BOUNCES: u32 = 4;
+pub(crate) const WORKGROUP: u32 = 8;
+
+#[repr(C)]
+#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
+pub(crate) struct DenoiseParams {
+    pub(crate) width: u32,
+    pub(crate) height: u32,
+    pub(crate) step: u32,
+    pub(crate) first: u32,
+    pub(crate) last: u32,
+    pub(crate) inv_sqrt_n: f32,
+    pub(crate) _pad: [u32; 2],
+}
+
+pub(crate) const DENOISE_ITERATIONS: usize = 3; // à-trous steps 1, 2, 4
+
+/// A scene as the tracer's buffers hold it.
+pub(crate) struct PackedScene {
+    pub tris: Vec<GpuTriangle>,
+    pub materials: Vec<GpuMaterial>,
+    /// Empty unless `with_bvh` was asked for (the ray-query tier builds
+    /// its own structure).
+    pub nodes: Vec<GpuBvhNode>,
+}
+
+/// Lay a scene out for the tracer. An image joins it as a quad of two
+/// triangles (corners in its own order, top-left first) under a material
+/// of its own, marked textured (`albedo.w`), past the scene's materials —
+/// or past the one stood in for a scene that names none — so every tier
+/// meets it as it meets any triangle, and a scene that is an image alone is
+/// not an empty one. `with_bvh` builds the BVH, reordering the triangles.
+pub(crate) fn pack_scene(
+    triangles: &[RtTriangle],
+    materials: &[RtMaterial],
+    image_corners: Option<[[f32; 3]; 4]>,
+    with_bvh: bool,
+) -> PackedScene {
+    let mut tris: Vec<RtTriangle> = triangles.to_vec();
+    let image_material = materials.len().max(1) as u32;
+    if let Some([tl, tr, br, bl]) = image_corners {
+        tris.push(RtTriangle { p0: tl, p1: bl, p2: tr, material: image_material });
+        tris.push(RtTriangle { p0: tr, p1: bl, p2: br, material: image_material });
+    }
+    let nodes = if with_bvh { build_bvh(&mut tris) } else { Vec::new() };
+    let gpu_tris = tris
+        .iter()
+        .map(|t| GpuTriangle {
+            p0: [t.p0[0], t.p0[1], t.p0[2], f32::from_bits(t.material)],
+            p1: [t.p1[0], t.p1[1], t.p1[2], 0.0],
+            p2: [t.p2[0], t.p2[1], t.p2[2], 0.0],
+        })
+        .collect();
+    let mut gpu_mats: Vec<GpuMaterial> = if materials.is_empty() {
+        vec![GpuMaterial { albedo: [0.8, 0.8, 0.8, 0.0], emission: [0.0; 4] }]
+    } else {
+        materials
+            .iter()
+            .map(|m| GpuMaterial {
+                albedo: [m.albedo[0], m.albedo[1], m.albedo[2], 0.0],
+                emission: [m.emission[0], m.emission[1], m.emission[2], 0.0],
+            })
+            .collect()
+    };
+    if image_corners.is_some() {
+        // albedo.w marks it textured: the shader takes the colour from the image.
+        gpu_mats.push(GpuMaterial { albedo: [1.0, 1.0, 1.0, 1.0], emission: [0.0; 4] });
+    }
+    PackedScene { tris: gpu_tris, materials: gpu_mats, nodes }
+}
+
+/// The image a frame traces, as its parameter block describes it: the
+/// texture's size and the quad's corners and opacity.
+pub(crate) struct ParamImage {
+    pub width: u32,
+    pub height: u32,
+    pub corners: [[f32; 3]; 4],
+    pub opacity: f32,
+}
+
+fn v4([x, y, z]: [f32; 3]) -> [f32; 4] {
+    [x, y, z, 0.0]
+}
+
+/// One dispatch's parameter block. With no image (or one not uploaded
+/// yet) the quad's opacity is 0, which lets every ray through it.
+#[allow(clippy::too_many_arguments)]
+pub(crate) fn rt_params(
+    camera: RtCamera,
+    size: (u32, u32),
+    sample_index: u32,
+    spp: u32,
+    image: Option<ParamImage>,
+    background: Option<[f32; 3]>,
+    environment: &RtEnvironment,
+) -> RtParams {
+    let (img_origin, img_u, img_v) = match image {
+        Some(ParamImage { width, height, corners: [tl, tr, _, bl], opacity }) => (
+            [tl[0], tl[1], tl[2], opacity.clamp(0.0, 1.0)],
+            [tr[0] - tl[0], tr[1] - tl[1], tr[2] - tl[2], width as f32],
+            [bl[0] - tl[0], bl[1] - tl[1], bl[2] - tl[2], height as f32],
+        ),
+        None => ([0.0; 4], [1.0, 0.0, 0.0, 1.0], [0.0, 1.0, 0.0, 1.0]),
+    };
+    RtParams {
+        inv_mvp: camera.inv_mvp,
+        width: size.0,
+        height: size.1,
+        sample_index,
+        max_bounces: MAX_BOUNCES,
+        spp,
+        _pad: [0; 3],
+        img_origin,
+        img_u,
+        img_v,
+        background: match background {
+            Some([r, g, b]) => [r, g, b, 1.0],
+            None => [0.0; 4],
+        },
+        sun_dir: v4(environment.sun_unit()),
+        sun_color: v4(environment.sun_color),
+        sky_zenith: v4(environment.sky_zenith),
+        sky_nadir: v4(environment.sky_nadir),
+    }
+}
+
+/// The denoiser's blocks, one per à-trous iteration (steps 1, 2, 4), for an
+/// accumulation that will hold `n_after` samples once this frame's dispatch
+/// lands — the colour sigma tightens as it grows.
+pub(crate) fn denoise_params(width: u32, height: u32, n_after: u32) -> [DenoiseParams; DENOISE_ITERATIONS] {
+    let inv_sqrt_n = 1.0 / (n_after.max(1) as f32).sqrt();
+    std::array::from_fn(|i| DenoiseParams {
+        width,
+        height,
+        step: 1 << i,
+        first: (i == 0) as u32,
+        last: (i == DENOISE_ITERATIONS - 1) as u32,
+        inv_sqrt_n,
+        _pad: [0; 2],
+    })
+}
+
+#[cfg(test)]
+mod tests {
+    use super::*;
+
+    // CPU mirror of the shader's traversal, for parity testing.
+    fn intersect_tri_cpu(ro: [f32; 3], rd: [f32; 3], t: &RtTriangle, t_limit: f32) -> f32 {
+        let sub = |a: [f32; 3], b: [f32; 3]| [a[0] - b[0], a[1] - b[1], a[2] - b[2]];
+        let cross = |a: [f32; 3], b: [f32; 3]| {
+            [
+                a[1] * b[2] - a[2] * b[1],
+                a[2] * b[0] - a[0] * b[2],
+                a[0] * b[1] - a[1] * b[0],
+            ]
+        };
+        let dot = |a: [f32; 3], b: [f32; 3]| a[0] * b[0] + a[1] * b[1] + a[2] * b[2];
+        let e1 = sub(t.p1, t.p0);
+        let e2 = sub(t.p2, t.p0);
+        let h = cross(rd, e2);
+        let a = dot(e1, h);
+        if a.abs() < 1e-8 {
+            return 1e30;
+        }
+        let f = 1.0 / a;
+        let s = sub(ro, t.p0);
+        let u = f * dot(s, h);
+        if !(0.0..=1.0).contains(&u) {
+            return 1e30;
+        }
+        let q = cross(s, e1);
+        let v = f * dot(rd, q);
+        if v < 0.0 || u + v > 1.0 {
+            return 1e30;
+        }
+        let tt = f * dot(e2, q);
+        if tt > 1e-4 && tt < t_limit {
+            return tt;
+        }
+        1e30
+    }
+
+    fn traverse_bvh_cpu(
+        nodes: &[GpuBvhNode],
+        tris: &[RtTriangle],
+        ro: [f32; 3],
+        rd: [f32; 3],
+    ) -> (f32, Option<usize>) {
+        if nodes.is_empty() {
+            return (1e30, None);
+        }
+        let inv = [1.0 / rd[0], 1.0 / rd[1], 1.0 / rd[2]];
+        let hit_aabb = |min: [f32; 3], max: [f32; 3], t_limit: f32| -> bool {
+            let mut tn = f32::NEG_INFINITY;
+            let mut tf = f32::INFINITY;
+            for a in 0..3 {
+                let t1 = (min[a] - ro[a]) * inv[a];
+                let t2 = (max[a] - ro[a]) * inv[a];
+                tn = tn.max(t1.min(t2));
+                tf = tf.min(t1.max(t2));
+            }
+            tf >= tn.max(0.0) && tn < t_limit
+        };
+        let mut best = 1e30f32;
+        let mut best_tri = None;
+        let mut stack = vec![0u32];
+        while let Some(idx) = stack.pop() {
+            let node = &nodes[idx as usize];
+            if !hit_aabb(node.min, node.max, best) {
+                continue;
+            }
+            if node.count > 0 {
+                for i in node.left_first..node.left_first + node.count {
+                    let t = intersect_tri_cpu(ro, rd, &tris[i as usize], best);
+                    if t < best {
+                        best = t;
+                        best_tri = Some(i as usize);
+                    }
+                }
+            } else {
+                stack.push(node.left_first);
+                stack.push(node.left_first + 1);
+            }
+        }
+        (best, best_tri)
+    }
+
+    fn brute_force(tris: &[RtTriangle], ro: [f32; 3], rd: [f32; 3]) -> (f32, Option<usize>) {
+        let mut best = 1e30f32;
+        let mut best_tri = None;
+        for (i, t) in tris.iter().enumerate() {
+            let tt = intersect_tri_cpu(ro, rd, t, best);
+            if tt < best {
+                best = tt;
+                best_tri = Some(i);
+            }
+        }
+        (best, best_tri)
+    }
+
+    // Deterministic LCG so the test needs no rand dependency.
+    struct Lcg(u64);
+    impl Lcg {
+        fn next_f32(&mut self) -> f32 {
+            self.0 = self.0.wrapping_mul(6364136223846793005).wrapping_add(1442695040888963407);
+            ((self.0 >> 33) as f32) / (u32::MAX >> 1) as f32
+        }
+        fn point(&mut self, scale: f32) -> [f32; 3] {
+            [
+                (self.next_f32() - 0.5) * scale,
+                (self.next_f32() - 0.5) * scale,
+                (self.next_f32() - 0.5) * scale,
+            ]
+        }
+    }
+
+    fn random_scene(n: usize, seed: u64) -> Vec<RtTriangle> {
+        let mut rng = Lcg(seed);
+        (0..n)
+            .map(|i| {
+                let c = rng.point(20.0);
+                let jitter = |rng: &mut Lcg, c: [f32; 3]| {
+                    let d = rng.point(2.0);
+                    [c[0] + d[0], c[1] + d[1], c[2] + d[2]]
+                };
+                RtTriangle {
+                    p0: jitter(&mut rng, c),
+                    p1: jitter(&mut rng, c),
+                    p2: jitter(&mut rng, c),
+                    material: (i % 5) as u32,
+                }
+            })
+            .collect()
+    }
+
+    #[test]
+    fn test_bvh_matches_brute_force() {
+        let mut tris = random_scene(500, 42);
+        let nodes = build_bvh(&mut tris);
+        assert!(!nodes.is_empty());
+        let mut rng = Lcg(7);
+        let mut hits = 0;
+        for _ in 0..200 {
+            let ro = rng.point(40.0);
+            let target = rng.point(10.0);
+            let d = [target[0] - ro[0], target[1] - ro[1], target[2] - ro[2]];
+            let len = (d[0] * d[0] + d[1] * d[1] + d[2] * d[2]).sqrt().max(1e-6);
+            let rd = [d[0] / len, d[1] / len, d[2] / len];
+            let (t_bvh, tri_bvh) = traverse_bvh_cpu(&nodes, &tris, ro, rd);
+            let (t_ref, tri_ref) = brute_force(&tris, ro, rd);
+            assert_eq!(tri_bvh, tri_ref, "different triangle hit");
+            assert!((t_bvh - t_ref).abs() < 1e-4, "t mismatch: {t_bvh} vs {t_ref}");
+            if tri_bvh.is_some() {
+                hits += 1;
+            }
+        }
+        assert!(hits > 20, "test rays barely hit the scene ({hits}/200)");
+    }
+
+    #[test]
+    fn test_bvh_leaf_ranges_cover_all_triangles() {
+        let mut tris = random_scene(300, 9);
+        let nodes = build_bvh(&mut tris);
+        let mut seen = vec![false; tris.len()];
+        for node in &nodes {
+            if node.count > 0 {
+                for i in node.left_first..node.left_first + node.count {
+                    assert!(!seen[i as usize], "triangle {i} in two leaves");
+                    seen[i as usize] = true;
+                }
+            }
+        }
+        assert!(seen.iter().all(|&s| s), "not every triangle is in a leaf");
+    }
+
+    #[test]
+    fn test_bvh_degenerate_identical_centroids() {
+        // All triangles share one centroid: SAH can't split, the median
+        // fallback must still terminate and cover everything.
+        let tri = RtTriangle {
+            p0: [0.0, 0.0, 0.0],
+            p1: [1.0, 0.0, 0.0],
+            p2: [0.0, 1.0, 0.0],
+            material: 0,
+        };
+        let mut tris = vec![tri; 100];
+        let nodes = build_bvh(&mut tris);
+        let covered: u32 = nodes.iter().filter(|n| n.count > 0).map(|n| n.count).sum();
+        assert_eq!(covered, 100);
+        let (t, hit) = traverse_bvh_cpu(&nodes, &tris, [0.2, 0.2, -5.0], [0.0, 0.0, 1.0]);
+        assert!(hit.is_some());
+        assert!((t - 5.0).abs() < 1e-3);
+    }
+
+    #[test]
+    fn test_bvh_empty_and_single() {
+        let mut empty: Vec<RtTriangle> = Vec::new();
+        assert!(build_bvh(&mut empty).is_empty());
+
+        let mut single = vec![RtTriangle {
+            p0: [-1.0, -1.0, 0.0],
+            p1: [1.0, -1.0, 0.0],
+            p2: [0.0, 1.0, 0.0],
+            material: 3,
+        }];
+        let nodes = build_bvh(&mut single);
+        assert_eq!(nodes.len(), 1);
+        assert_eq!(nodes[0].count, 1);
+        let (t, hit) = traverse_bvh_cpu(&nodes, &single, [0.0, 0.0, -3.0], [0.0, 0.0, 1.0]);
+        assert_eq!(hit, Some(0));
+        assert!((t - 3.0).abs() < 1e-4);
+    }
+}
diff --git a/src/vk/rt_bvh.wgsl b/src/draw/rt_bvh.wgsl
similarity index 100%
rename from src/vk/rt_bvh.wgsl
rename to src/draw/rt_bvh.wgsl
diff --git a/src/vk/rt_common.wgsl b/src/draw/rt_common.wgsl
similarity index 100%
rename from src/vk/rt_common.wgsl
rename to src/draw/rt_common.wgsl
diff --git a/src/vk/rt_denoise.wgsl b/src/draw/rt_denoise.wgsl
similarity index 100%
rename from src/vk/rt_denoise.wgsl
rename to src/draw/rt_denoise.wgsl
diff --git a/src/vk/rt_query.wgsl b/src/draw/rt_query.wgsl
similarity index 100%
rename from src/vk/rt_query.wgsl
rename to src/draw/rt_query.wgsl
diff --git a/src/draw/scene.rs b/src/draw/scene.rs
index 9124662..7e7ff1a 100644
--- a/src/draw/scene.rs
+++ b/src/draw/scene.rs
@@ -8,7 +8,9 @@
 //! image, scissored to a pane, which the renderer copies beneath the UI pass
 //! and the 2D shader's blur plates sample. A staged scene is drawn once; the
 //! backdrop it leaves is shown under every later frame until the next one.
-//! `scene3d.wgsl` / `scene3d_image.wgsl` (in `draw/`) are the shaders.
+//! `scene3d.wgsl` / `scene3d_image.wgsl` (in `draw/`) are the shaders. A
+//! traced pane ([`super::rt`]) takes the raster scene's place in the same
+//! backdrop, staged through the same trait.
 
 /// Layout-identical to the app's `geometry::Vertex3D` (bytemuck-castable at cutover).
 #[repr(C)]
@@ -128,6 +130,8 @@ pub(crate) struct SceneUniforms {
 /// it with its INWARD derivative normal, said the right way round.
 pub(crate) const DEFAULT_SCENE_LIGHT: [f32; 3] = [0.55, -0.45, -0.7];
 
+use super::rt::{RtCamera, RtEnvironment, RtImage, RtMaterial, RtTriangle};
+
 /// What an app stages a 3D scene through: the renderer's half of the
 /// pass, the same on every renderer (`vk::VkRenderer`, `web::WebRenderer`).
 /// An app takes one in `Application::init_3d` (make its meshes) and
@@ -146,6 +150,28 @@ pub trait Stage3D {
     /// The direction TOWARD the flat shading's light, in world space; set
     /// it once, and light a `prelit` mesh by the same vector.
     fn set_scene_light(&mut self, toward: [f32; 3]);
+
+    /// Replace the path tracer's scene (triangles in the space the camera's
+    /// `inv_mvp` unprojects into); the BVH is built on the CPU. Rare: a
+    /// geometry rebuild. Restarts the accumulation.
+    fn set_rt_scene(&mut self, triangles: &[RtTriangle], materials: &[RtMaterial]) {
+        self.set_rt_scene_with_image(triangles, materials, None);
+    }
+    /// [`set_rt_scene`](Self::set_rt_scene) with an uploaded image standing
+    /// in the scene (the picture the raster pass draws as a `SceneImage`).
+    fn set_rt_scene_with_image(&mut self, triangles: &[RtTriangle], materials: &[RtMaterial], image: Option<RtImage>);
+    /// The traced scene's sky and sun. A change restarts the accumulation.
+    fn set_rt_environment(&mut self, environment: RtEnvironment);
+    /// What a camera ray that meets nothing shows (linear RGB), or `None`
+    /// for the sky. A change restarts the accumulation.
+    fn set_rt_background(&mut self, color: Option<[f32; 3]>);
+    /// Stage one progressive pass into `pane` (physical px) for the next
+    /// frame, in place of the raster scene there. Call it every frame while
+    /// tracing: each adds a sample; a camera, pane or scene change restarts.
+    fn stage_rt(&mut self, pane: (u32, u32, u32, u32), camera: RtCamera);
+    /// True while another staged frame would still refine the traced image
+    /// — the app's cue to keep asking for frames.
+    fn rt_accumulating(&self) -> bool;
 }
 
 /// A staged scene's uniform blocks, in the order its draws use them: one per
diff --git a/src/draw/shaders.rs b/src/draw/shaders.rs
index 4568d0c..8ad7000 100644
--- a/src/draw/shaders.rs
+++ b/src/draw/shaders.rs
@@ -20,6 +20,19 @@ pub const SCENE3D: &str = include_str!("scene3d.wgsl");
 /// The 3D scene pass's images: textured quads under the meshes' uniforms.
 pub const SCENE3D_IMAGE: &str = include_str!("scene3d_image.wgsl");
 
+/// The path tracer ([`super::rt`]): its shared core, then one trace tier
+/// after it — the compute BVH traversal, or (Vulkan only) hardware ray
+/// queries — and the à-trous denoiser that runs over its output.
+pub const RT_COMMON: &str = include_str!("rt_common.wgsl");
+pub const RT_BVH: &str = include_str!("rt_bvh.wgsl");
+pub const RT_QUERY: &str = include_str!("rt_query.wgsl");
+pub const RT_DENOISE: &str = include_str!("rt_denoise.wgsl");
+
+/// The compute tier's tracer: [`RT_COMMON`] with [`RT_BVH`] after it.
+pub fn rt_bvh_source() -> String {
+    format!("{RT_COMMON}\n{RT_BVH}")
+}
+
 /// The one line that differs between the Vulkan and the WebGPU 2D shader.
 const PUSH_BLOCK: &str = "var<push_constant> rrect_clip: RRectClip;";
 const UNIFORM_BLOCK: &str = "@group(1) @binding(0) var<uniform> rrect_clip: RRectClip;";
diff --git a/src/engine.rs b/src/engine.rs
index 09c9092..6669e62 100644
--- a/src/engine.rs
+++ b/src/engine.rs
@@ -5,6 +5,7 @@
 pub use crate::backend::app::{
     Application, AppSender, LogicalPosition, LogicalSize, RenderContext, Stage3D, WindowAction, WindowSettings,
 };
+pub use crate::draw::rt::{RtCamera, RtEnvironment, RtImage, RtMaterial, RtTriangle};
 pub use crate::draw::scene::{MeshId, SceneDraw, SceneImage, Vertex3D};
 pub use crate::backend::driver::PressedKey;
 pub use crate::backend::tessellate::{
diff --git a/src/vk/renderer.rs b/src/vk/renderer.rs
index d032430..8e0768d 100644
--- a/src/vk/renderer.rs
+++ b/src/vk/renderer.rs
@@ -2346,6 +2346,21 @@ impl crate::draw::scene::Stage3D for VkRenderer {
     fn set_scene_light(&mut self, toward: [f32; 3]) {
         VkRenderer::set_scene_light(self, toward)
     }
+    fn set_rt_scene_with_image(&mut self, triangles: &[RtTriangle], materials: &[RtMaterial], image: Option<RtImage>) {
+        VkRenderer::set_rt_scene_with_image(self, triangles, materials, image)
+    }
+    fn set_rt_environment(&mut self, environment: RtEnvironment) {
+        VkRenderer::set_rt_environment(self, environment)
+    }
+    fn set_rt_background(&mut self, color: Option<[f32; 3]>) {
+        VkRenderer::set_rt_background(self, color)
+    }
+    fn stage_rt(&mut self, pane: (u32, u32, u32, u32), camera: RtCamera) {
+        VkRenderer::stage_rt(self, pane, camera)
+    }
+    fn rt_accumulating(&self) -> bool {
+        VkRenderer::rt_accumulating(self)
+    }
 }
 
 impl Drop for VkRenderer {
diff --git a/src/vk/rt.rs b/src/vk/rt.rs
index 361162b..540c597 100644
--- a/src/vk/rt.rs
+++ b/src/vk/rt.rs
@@ -19,54 +19,11 @@ use gpu_allocator::MemoryLocation;
 use super::renderer::{
     compile_wgsl, compile_wgsl_ray_query, create_cpu_buffer, destroy_cpu_buffer, AllocatedBuffer,
 };
-
-/// One triangle of an RT scene, in the same space as the camera's `inv_mvp`
-/// (for the designer: mesh space, the space `Vertex3D` positions live in).
-#[derive(Debug, Clone, Copy, PartialEq)]
-pub struct RtTriangle {
-    pub p0: [f32; 3],
-    pub p1: [f32; 3],
-    pub p2: [f32; 3],
-    /// Index into the material slice passed alongside.
-    pub material: u32,
-}
-
-/// Lambertian surface + optional emission, linear color (matching the raster
-/// path, whose vertex colors land in the sRGB attachment as linear values).
-#[derive(Debug, Clone, Copy, PartialEq)]
-pub struct RtMaterial {
-    pub albedo: [f32; 3],
-    pub emission: [f32; 3],
-}
-
-/// An image standing in the traced scene: a quad that shows the image's
-/// colour as it is — unlit, as the raster pass's `SceneImage` draws it — and
-/// lets a ray through where the image is clear. A path ends on it, so to
-/// the rest of the scene it is a light of its own colour. The tracer adds
-/// the quad to the scene itself.
-///
-/// The image is one the 2D pass already holds (an id from `upload_rgba`),
-/// so a picture shown in the raster viewport costs the tracer nothing more.
-#[derive(Debug, Clone, Copy, PartialEq)]
-pub struct RtImage {
-    pub image: u32,
-    /// The quad's corners in the scene's space, in the image's own order:
-    /// top-left, top-right, bottom-right, bottom-left.
-    pub corners: [[f32; 3]; 4],
-    /// Alpha multiplier over the image's own.
-    pub opacity: f32,
-}
-
-/// [`RtImage`] for the headless tracer, which has no 2D pass to share an
-/// image with and is handed the pixels: tightly packed sRGB RGBA8.
-#[derive(Debug, Clone, Copy)]
-pub struct RtImagePixels<'a> {
-    pub pixels: &'a [u8],
-    pub width: u32,
-    pub height: u32,
-    pub corners: [[f32; 3]; 4],
-    pub opacity: f32,
-}
+pub use crate::draw::rt::{RtCamera, RtEnvironment, RtImage, RtImagePixels, RtMaterial, RtTriangle};
+use crate::draw::rt::{
+    denoise_params, pack_scene, rt_params, DenoiseParams, ParamImage, RtParams, DENOISE_ITERATIONS, MAX_SAMPLES,
+    WORKGROUP,
+};
 
 /// Where the stage's image comes from.
 pub(crate) enum RtImageSource<'a> {
@@ -77,285 +34,6 @@ pub(crate) enum RtImageSource<'a> {
     Pixels { pixels: &'a [u8], width: u32, height: u32 },
 }
 
-/// The full camera: the inverse of the raster path's `proj * view * model`.
-/// Rays are unprojected from NDC through it, so any matrix stack that renders
-/// the raster viewport drives the tracer unchanged.
-#[derive(Debug, Clone, Copy, PartialEq)]
-pub struct RtCamera {
-    pub inv_mvp: [[f32; 4]; 4],
-}
-
-// --- GPU layouts (must match rt.wgsl) ---
-
-#[repr(C)]
-#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
-struct GpuTriangle {
-    p0: [f32; 4], // w = material index (bitcast)
-    p1: [f32; 4],
-    p2: [f32; 4],
-}
-
-#[repr(C)]
-#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
-struct GpuMaterial {
-    albedo: [f32; 4],
-    emission: [f32; 4],
-}
-
-#[repr(C)]
-#[derive(Debug, Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
-pub(crate) struct GpuBvhNode {
-    pub(crate) min: [f32; 3],
-    /// Leaf (`count > 0`): first triangle. Internal: left child; right = +1.
-    pub(crate) left_first: u32,
-    pub(crate) max: [f32; 3],
-    pub(crate) count: u32,
-}
-
-#[repr(C)]
-#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
-struct RtParams {
-    inv_mvp: [[f32; 4]; 4],
-    width: u32,
-    height: u32,
-    sample_index: u32,
-    max_bounces: u32,
-    spp: u32,
-    _pad: [u32; 3],
-    img_origin: [f32; 4],
-    img_u: [f32; 4],
-    img_v: [f32; 4],
-    background: [f32; 4],
-    // The environment, each xyz in a vec4 (w unused): toward the sun
-    // (unit), the sun's radiance, the sky overhead, the sky below.
-    sun_dir: [f32; 4],
-    sun_color: [f32; 4],
-    sky_zenith: [f32; 4],
-    sky_nadir: [f32; 4],
-}
-
-/// The traced scene's light: a sky that grades from `sky_nadir` straight
-/// down to `sky_zenith` straight up, and a sun — a bright lobe toward
-/// `sun_direction` of `sun_color` radiance. It is the tracer's ONLY light
-/// (nothing in a scene emits unless its material does). Colours are linear
-/// RGB and may exceed 1. [`Default`] is the studio sky the tracer always
-/// had.
-#[derive(Clone, Copy, Debug, PartialEq)]
-pub struct RtEnvironment {
-    /// Toward the sun, world space; any length (zero is straight up).
-    pub sun_direction: [f32; 3],
-    pub sun_color: [f32; 3],
-    pub sky_zenith: [f32; 3],
-    pub sky_nadir: [f32; 3],
-}
-
-impl Default for RtEnvironment {
-    fn default() -> Self {
-        Self {
-            sun_direction: [0.45, 0.75, 0.35],
-            sun_color: [8.0, 7.6, 6.8],
-            sky_zenith: [0.72, 0.82, 0.98],
-            sky_nadir: [0.32, 0.31, 0.35],
-        }
-    }
-}
-
-impl RtEnvironment {
-    fn sun_unit(&self) -> [f32; 3] {
-        let v = glam::Vec3::from_array(self.sun_direction);
-        if v.length_squared() > 1e-12 { v.normalize().to_array() } else { [0.0, 1.0, 0.0] }
-    }
-}
-
-// --- BVH construction (binned SAH) ---
-
-const BVH_BINS: usize = 8;
-const BVH_LEAF_MAX: u32 = 4;
-
-#[derive(Clone, Copy)]
-struct Aabb {
-    min: [f32; 3],
-    max: [f32; 3],
-}
-
-impl Aabb {
-    const EMPTY: Aabb = Aabb { min: [f32::INFINITY; 3], max: [f32::NEG_INFINITY; 3] };
-
-    fn grow(&mut self, p: [f32; 3]) {
-        for a in 0..3 {
-            self.min[a] = self.min[a].min(p[a]);
-            self.max[a] = self.max[a].max(p[a]);
-        }
-    }
-
-    fn grow_aabb(&mut self, other: &Aabb) {
-        self.grow(other.min);
-        self.grow(other.max);
-    }
-
-    fn half_area(&self) -> f32 {
-        let dx = (self.max[0] - self.min[0]).max(0.0);
-        let dy = (self.max[1] - self.min[1]).max(0.0);
-        let dz = (self.max[2] - self.min[2]).max(0.0);
-        dx * dy + dy * dz + dz * dx
-    }
-}
-
-fn tri_aabb(t: &RtTriangle) -> Aabb {
-    let mut b = Aabb::EMPTY;
-    b.grow(t.p0);
-    b.grow(t.p1);
-    b.grow(t.p2);
-    b
-}
-
-fn tri_centroid(t: &RtTriangle) -> [f32; 3] {
-    let mut c = [0.0f32; 3];
-    for a in 0..3 {
-        c[a] = (t.p0[a] + t.p1[a] + t.p2[a]) / 3.0;
-    }
-    c
-}
-
-/// Build a BVH over `triangles`, reordering them so leaves reference
-/// contiguous ranges. Returns the flat node array (empty input → empty vec).
-pub(crate) fn build_bvh(triangles: &mut Vec<RtTriangle>) -> Vec<GpuBvhNode> {
-    if triangles.is_empty() {
-        return Vec::new();
-    }
-    let bounds: Vec<Aabb> = triangles.iter().map(tri_aabb).collect();
-    let centroids: Vec<[f32; 3]> = triangles.iter().map(tri_centroid).collect();
-    let mut order: Vec<u32> = (0..triangles.len() as u32).collect();
-
-    fn range_bounds(order: &[u32], bounds: &[Aabb]) -> Aabb {
-        let mut b = Aabb::EMPTY;
-        for &i in order {
-            b.grow_aabb(&bounds[i as usize]);
-        }
-        b
-    }
-
-    let mut nodes: Vec<GpuBvhNode> = Vec::with_capacity(triangles.len() * 2);
-    let root_bounds = range_bounds(&order, &bounds);
-    nodes.push(GpuBvhNode {
-        min: root_bounds.min,
-        left_first: 0,
-        max: root_bounds.max,
-        count: triangles.len() as u32,
-    });
-
-    // (node index, start, count) work list over `order`.
-    let mut work = vec![(0usize, 0usize, triangles.len())];
-    while let Some((node_idx, start, count)) = work.pop() {
-        if (count as u32) <= BVH_LEAF_MAX {
-            continue; // stays a leaf
-        }
-        let slice = &mut order[start..start + count];
-
-        // Centroid bounds pick the split axis.
-        let mut cb = Aabb::EMPTY;
-        for &i in slice.iter() {
-            cb.grow(centroids[i as usize]);
-        }
-        let mut axis = 0;
-        let mut extent = 0.0f32;
-        for a in 0..3 {
-            let e = cb.max[a] - cb.min[a];
-            if e > extent {
-                extent = e;
-                axis = a;
-            }
-        }
-
-        let mut split_at = None;
-        if extent > 1e-12 {
-            // Binned SAH along `axis`.
-            let scale = BVH_BINS as f32 / extent;
-            let bin_of = |i: u32| -> usize {
-                (((centroids[i as usize][axis] - cb.min[axis]) * scale) as usize)
-                    .min(BVH_BINS - 1)
-            };
-            let mut bin_bounds = [Aabb::EMPTY; BVH_BINS];
-            let mut bin_counts = [0usize; BVH_BINS];
-            for &i in slice.iter() {
-                let b = bin_of(i);
-                bin_counts[b] += 1;
-                bin_bounds[b].grow_aabb(&bounds[i as usize]);
-            }
-            // Cost of each of the BINS-1 split planes.
-            let mut best_cost = f32::INFINITY;
-            let mut best_plane = 0usize;
-            for plane in 1..BVH_BINS {
-                let (mut lb, mut rb) = (Aabb::EMPTY, Aabb::EMPTY);
-                let (mut lc, mut rc) = (0usize, 0usize);
-                for b in 0..plane {
-                    lb.grow_aabb(&bin_bounds[b]);
-                    lc += bin_counts[b];
-                }
-                for b in plane..BVH_BINS {
-                    rb.grow_aabb(&bin_bounds[b]);
-                    rc += bin_counts[b];
-                }
-                if lc == 0 || rc == 0 {
-                    continue;
-                }
-                let cost = lb.half_area() * lc as f32 + rb.half_area() * rc as f32;
-                if cost < best_cost {
-                    best_cost = cost;
-                    best_plane = plane;
-                }
-            }
-            if best_plane > 0 {
-                let mut mid = 0usize;
-                for k in 0..count {
-                    if bin_of(slice[k]) < best_plane {
-                        slice.swap(k, mid);
-                        mid += 1;
-                    }
-                }
-                if mid > 0 && mid < count {
-                    split_at = Some(mid);
-                }
-            }
-        }
-        // Degenerate centroids or a one-sided SAH result: median split keeps
-        // the tree balanced instead of forcing a giant leaf.
-        let mid = split_at.unwrap_or(count / 2);
-
-        let left_bounds = range_bounds(&slice[..mid], &bounds);
-        let right_bounds = range_bounds(&slice[mid..], &bounds);
-        let left_idx = nodes.len();
-        nodes.push(GpuBvhNode {
-            min: left_bounds.min,
-            left_first: (start) as u32,
-            max: left_bounds.max,
-            count: mid as u32,
-        });
-        nodes.push(GpuBvhNode {
-            min: right_bounds.min,
-            left_first: (start + mid) as u32,
-            max: right_bounds.max,
-            count: (count - mid) as u32,
-        });
-        nodes[node_idx].left_first = left_idx as u32;
-        nodes[node_idx].count = 0;
-        work.push((left_idx, start, mid));
-        work.push((left_idx + 1, start + mid, count - mid));
-    }
-
-    // Apply the final order to the triangle array so leaf ranges are direct.
-    let reordered: Vec<RtTriangle> =
-        order.iter().map(|&i| triangles[i as usize]).collect();
-    *triangles = reordered;
-    nodes
-}
-
-// --- The Vulkan stage ---
-
-const MAX_SAMPLES: u32 = 1024;
-const MAX_BOUNCES: u32 = 4;
-const WORKGROUP: u32 = 8;
-
 /// The two trace backends. They share every shader line except
 /// `intersect_scene` (rt_bvh.wgsl vs rt_query.wgsl) and binding 1.
 #[derive(Debug, Clone, Copy, PartialEq, Eq)]
@@ -560,20 +238,6 @@ struct StagedImage {
     opacity: f32,
 }
 
-#[repr(C)]
-#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
-struct DenoiseParams {
-    width: u32,
-    height: u32,
-    step: u32,
-    first: u32,
-    last: u32,
-    inv_sqrt_n: f32,
-    _pad: [u32; 2],
-}
-
-const DENOISE_ITERATIONS: usize = 3; // à-trous steps 1, 2, 4
-
 struct DenoiserFrame {
     /// DENOISE_ITERATIONS dynamic-offset slices of [`DenoiseParams`].
     uniforms: AllocatedBuffer,
@@ -649,7 +313,7 @@ impl Denoiser {
                     None,
                 )
                 .expect("Failed to create denoise pipeline layout");
-            let spirv = compile_wgsl(include_str!("rt_denoise.wgsl"));
+            let spirv = compile_wgsl(crate::draw::shaders::RT_DENOISE);
             let shader_module = device
                 .create_shader_module(&vk::ShaderModuleCreateInfo::default().code(&spirv), None)
                 .expect("Failed to create denoise shader module");
@@ -790,22 +454,12 @@ impl Denoiser {
     /// sample count the accumulation will hold once this frame's dispatch
     /// lands — the color sigma tightens as it grows.
     fn write_frame_uniforms(&mut self, frame_index: usize, width: u32, height: u32, n_after: u32) {
-        let inv_sqrt_n = 1.0 / (n_after.max(1) as f32).sqrt();
         let frame = &mut self.frames[frame_index];
         let mapped = frame.uniforms.allocation.as_mut().unwrap().mapped_slice_mut().unwrap();
-        for i in 0..DENOISE_ITERATIONS {
-            let params = DenoiseParams {
-                width,
-                height,
-                step: 1 << i,
-                first: (i == 0) as u32,
-                last: (i == DENOISE_ITERATIONS - 1) as u32,
-                inv_sqrt_n,
-                _pad: [0; 2],
-            };
+        for (i, params) in denoise_params(width, height, n_after).iter().enumerate() {
             let offset = self.uniform_stride as usize * i;
             mapped[offset..offset + std::mem::size_of::<DenoiseParams>()]
-                .copy_from_slice(bytemuck::bytes_of(&params));
+                .copy_from_slice(bytemuck::bytes_of(params));
         }
     }
 
@@ -914,10 +568,6 @@ pub(crate) struct RtStage {
     environment: RtEnvironment,
 }
 
-fn v4([x, y, z]: [f32; 3]) -> [f32; 4] {
-    [x, y, z, 0.0]
-}
-
 impl RtStage {
     /// `accel_loader` present means the device has the ray-query stack; the
     /// stage then runs tier 2 unless `CCE_VK_RT=compute` forces the BVH tier.
@@ -1013,16 +663,8 @@ impl RtStage {
                 .expect("Failed to create RT pipeline layout");
 
             let spirv = match tier {
-                RtTier::Compute => compile_wgsl(&format!(
-                    "{}\n{}",
-                    include_str!("rt_common.wgsl"),
-                    include_str!("rt_bvh.wgsl")
-                )),
-                RtTier::RayQuery => compile_wgsl_ray_query(&format!(
-                    "{}\n{}",
-                    include_str!("rt_common.wgsl"),
-                    include_str!("rt_query.wgsl")
-                )),
+                RtTier::Compute => compile_wgsl(&crate::draw::shaders::rt_bvh_source()),
+                RtTier::RayQuery => compile_wgsl_ray_query(&format!("{}\n{}", crate::draw::shaders::RT_COMMON, crate::draw::shaders::RT_QUERY)),
             };
             let shader_module = device
                 .create_shader_module(&vk::ShaderModuleCreateInfo::default().code(&spirv), None)
@@ -1249,17 +891,10 @@ impl RtStage {
         materials: &[RtMaterial],
         image: Option<(RtImageSource, [[f32; 3]; 4], f32)>,
     ) {
-        let mut tris: Vec<RtTriangle> = triangles.to_vec();
-        // The index the image's material will have, past the scene's own
-        // (or past the one stood in for a scene that names none).
-        let image_material = materials.len().max(1) as u32;
-        if let Some((_, [tl, tr, br, bl], _)) = &image {
-            tris.push(RtTriangle { p0: *tl, p1: *bl, p2: *tr, material: image_material });
-            tris.push(RtTriangle { p0: *tr, p1: *bl, p2: *br, material: image_material });
-        }
         if let Some(mut old) = self.image.take().and_then(|i| i.owned) {
             old.destroy(device, allocator);
         }
+        let corners = image.as_ref().map(|(_, corners, _)| *corners);
         self.image = image.map(|(source, corners, opacity)| match source {
             RtImageSource::Shared(id) => {
                 StagedImage { shared: Some(id), owned: None, corners, opacity }
@@ -1279,34 +914,8 @@ impl RtStage {
                 opacity,
             },
         });
-        let nodes = match self.tier {
-            RtTier::Compute => build_bvh(&mut tris),
-            RtTier::RayQuery => Vec::new(),
-        };
-        let gpu_tris: Vec<GpuTriangle> = tris
-            .iter()
-            .map(|t| GpuTriangle {
-                p0: [t.p0[0], t.p0[1], t.p0[2], f32::from_bits(t.material)],
-                p1: [t.p1[0], t.p1[1], t.p1[2], 0.0],
-                p2: [t.p2[0], t.p2[1], t.p2[2], 0.0],
-            })
-            .collect();
-        let mut gpu_mats: Vec<GpuMaterial> = if materials.is_empty() {
-            vec![GpuMaterial { albedo: [0.8, 0.8, 0.8, 0.0], emission: [0.0; 4] }]
-        } else {
-            materials
-                .iter()
-                .map(|m| GpuMaterial {
-                    albedo: [m.albedo[0], m.albedo[1], m.albedo[2], 0.0],
-                    emission: [m.emission[0], m.emission[1], m.emission[2], 0.0],
-                })
-                .collect()
-        };
-        if self.image.is_some() {
-            // albedo.w marks it textured: the shader takes the colour from
-            // the image.
-            gpu_mats.push(GpuMaterial { albedo: [1.0, 1.0, 1.0, 1.0], emission: [0.0; 4] });
-        }
+        let packed = pack_scene(triangles, materials, corners, self.tier == RtTier::Compute);
+        let (gpu_tris, gpu_mats, nodes) = (packed.tris, packed.materials, packed.nodes);
 
         self.destroy_accel(device, allocator);
         for buf in [&mut self.nodes, &mut self.tris, &mut self.materials] {
@@ -1347,7 +956,7 @@ impl RtStage {
             vk::BufferUsageFlags::STORAGE_BUFFER,
             "rt-materials",
         );
-        self.tri_count = tris.len() as u32;
+        self.tri_count = gpu_tris.len() as u32;
         self.sample_index = 0;
 
         // Binding 1 (per tier), then the shared 2/3.
@@ -1884,18 +1493,13 @@ impl RtStage {
                 (None, Some(id)) => shared(id)?,
                 (None, None) => return None,
             };
-            Some((view, w, h, image.corners, image.opacity))
+            Some((view, ParamImage { width: w, height: h, corners: image.corners, opacity: image.opacity }))
         });
-        let (view, img_origin, img_u, img_v) = match bound {
-            Some((view, w, h, [tl, tr, _, bl], opacity)) => (
-                view,
-                [tl[0], tl[1], tl[2], opacity.clamp(0.0, 1.0)],
-                [tr[0] - tl[0], tr[1] - tl[1], tr[2] - tl[2], w as f32],
-                [bl[0] - tl[0], bl[1] - tl[1], bl[2] - tl[2], h as f32],
-            ),
-            // No image, or one not there yet: opacity 0 lets every ray
-            // through the quad, and the stand-in is never sampled for it.
-            None => (self.stand_in.view, [0.0; 4], [1.0, 0.0, 0.0, 1.0], [0.0, 1.0, 0.0, 1.0]),
+        // No image, or one not there yet: opacity 0 lets every ray through
+        // the quad, and the stand-in is never sampled for it.
+        let (view, param_image) = match bound {
+            Some((view, p)) => (view, Some(p)),
+            None => (self.stand_in.view, None),
         };
         Self::write_image_descriptor(
             device,
@@ -1903,26 +1507,15 @@ impl RtStage {
             view,
             self.image_sampler,
         );
-        let params = RtParams {
-            inv_mvp: camera.inv_mvp,
-            width: self.output_size.0,
-            height: self.output_size.1,
-            sample_index: self.sample_index,
-            max_bounces: MAX_BOUNCES,
-            spp: self.spp,
-            _pad: [0; 3],
-            img_origin,
-            img_u,
-            img_v,
-            background: match self.background {
-                Some([r, g, b]) => [r, g, b, 1.0],
-                None => [0.0; 4],
-            },
-            sun_dir: v4(self.environment.sun_unit()),
-            sun_color: v4(self.environment.sun_color),
-            sky_zenith: v4(self.environment.sky_zenith),
-            sky_nadir: v4(self.environment.sky_nadir),
-        };
+        let params = rt_params(
+            camera,
+            self.output_size,
+            self.sample_index,
+            self.spp,
+            param_image,
+            self.background,
+            &self.environment,
+        );
         let frame = &mut self.frames[frame_index];
         frame.uniforms.allocation.as_mut().unwrap().mapped_slice_mut().unwrap()
             [..std::mem::size_of::<RtParams>()]
@@ -2512,229 +2105,14 @@ impl Drop for RtOffscreen {
 mod tests {
     use super::*;
 
-    // CPU mirror of the shader's traversal, for parity testing.
-    fn intersect_tri_cpu(ro: [f32; 3], rd: [f32; 3], t: &RtTriangle, t_limit: f32) -> f32 {
-        let sub = |a: [f32; 3], b: [f32; 3]| [a[0] - b[0], a[1] - b[1], a[2] - b[2]];
-        let cross = |a: [f32; 3], b: [f32; 3]| {
-            [
-                a[1] * b[2] - a[2] * b[1],
-                a[2] * b[0] - a[0] * b[2],
-                a[0] * b[1] - a[1] * b[0],
-            ]
-        };
-        let dot = |a: [f32; 3], b: [f32; 3]| a[0] * b[0] + a[1] * b[1] + a[2] * b[2];
-        let e1 = sub(t.p1, t.p0);
-        let e2 = sub(t.p2, t.p0);
-        let h = cross(rd, e2);
-        let a = dot(e1, h);
-        if a.abs() < 1e-8 {
-            return 1e30;
-        }
-        let f = 1.0 / a;
-        let s = sub(ro, t.p0);
-        let u = f * dot(s, h);
-        if !(0.0..=1.0).contains(&u) {
-            return 1e30;
-        }
-        let q = cross(s, e1);
-        let v = f * dot(rd, q);
-        if v < 0.0 || u + v > 1.0 {
-            return 1e30;
-        }
-        let tt = f * dot(e2, q);
-        if tt > 1e-4 && tt < t_limit {
-            return tt;
-        }
-        1e30
-    }
-
-    fn traverse_bvh_cpu(
-        nodes: &[GpuBvhNode],
-        tris: &[RtTriangle],
-        ro: [f32; 3],
-        rd: [f32; 3],
-    ) -> (f32, Option<usize>) {
-        if nodes.is_empty() {
-            return (1e30, None);
-        }
-        let inv = [1.0 / rd[0], 1.0 / rd[1], 1.0 / rd[2]];
-        let hit_aabb = |min: [f32; 3], max: [f32; 3], t_limit: f32| -> bool {
-            let mut tn = f32::NEG_INFINITY;
-            let mut tf = f32::INFINITY;
-            for a in 0..3 {
-                let t1 = (min[a] - ro[a]) * inv[a];
-                let t2 = (max[a] - ro[a]) * inv[a];
-                tn = tn.max(t1.min(t2));
-                tf = tf.min(t1.max(t2));
-            }
-            tf >= tn.max(0.0) && tn < t_limit
-        };
-        let mut best = 1e30f32;
-        let mut best_tri = None;
-        let mut stack = vec![0u32];
-        while let Some(idx) = stack.pop() {
-            let node = &nodes[idx as usize];
-            if !hit_aabb(node.min, node.max, best) {
-                continue;
-            }
-            if node.count > 0 {
-                for i in node.left_first..node.left_first + node.count {
-                    let t = intersect_tri_cpu(ro, rd, &tris[i as usize], best);
-                    if t < best {
-                        best = t;
-                        best_tri = Some(i as usize);
-                    }
-                }
-            } else {
-                stack.push(node.left_first);
-                stack.push(node.left_first + 1);
-            }
-        }
-        (best, best_tri)
-    }
-
-    fn brute_force(tris: &[RtTriangle], ro: [f32; 3], rd: [f32; 3]) -> (f32, Option<usize>) {
-        let mut best = 1e30f32;
-        let mut best_tri = None;
-        for (i, t) in tris.iter().enumerate() {
-            let tt = intersect_tri_cpu(ro, rd, t, best);
-            if tt < best {
-                best = tt;
-                best_tri = Some(i);
-            }
-        }
-        (best, best_tri)
-    }
-
-    // Deterministic LCG so the test needs no rand dependency.
-    struct Lcg(u64);
-    impl Lcg {
-        fn next_f32(&mut self) -> f32 {
-            self.0 = self.0.wrapping_mul(6364136223846793005).wrapping_add(1442695040888963407);
-            ((self.0 >> 33) as f32) / (u32::MAX >> 1) as f32
-        }
-        fn point(&mut self, scale: f32) -> [f32; 3] {
-            [
-                (self.next_f32() - 0.5) * scale,
-                (self.next_f32() - 0.5) * scale,
-                (self.next_f32() - 0.5) * scale,
-            ]
-        }
-    }
-
-    fn random_scene(n: usize, seed: u64) -> Vec<RtTriangle> {
-        let mut rng = Lcg(seed);
-        (0..n)
-            .map(|i| {
-                let c = rng.point(20.0);
-                let jitter = |rng: &mut Lcg, c: [f32; 3]| {
-                    let d = rng.point(2.0);
-                    [c[0] + d[0], c[1] + d[1], c[2] + d[2]]
-                };
-                RtTriangle {
-                    p0: jitter(&mut rng, c),
-                    p1: jitter(&mut rng, c),
-                    p2: jitter(&mut rng, c),
-                    material: (i % 5) as u32,
-                }
-            })
-            .collect()
-    }
-
-    #[test]
-    fn test_bvh_matches_brute_force() {
-        let mut tris = random_scene(500, 42);
-        let nodes = build_bvh(&mut tris);
-        assert!(!nodes.is_empty());
-        let mut rng = Lcg(7);
-        let mut hits = 0;
-        for _ in 0..200 {
-            let ro = rng.point(40.0);
-            let target = rng.point(10.0);
-            let d = [target[0] - ro[0], target[1] - ro[1], target[2] - ro[2]];
-            let len = (d[0] * d[0] + d[1] * d[1] + d[2] * d[2]).sqrt().max(1e-6);
-            let rd = [d[0] / len, d[1] / len, d[2] / len];
-            let (t_bvh, tri_bvh) = traverse_bvh_cpu(&nodes, &tris, ro, rd);
-            let (t_ref, tri_ref) = brute_force(&tris, ro, rd);
-            assert_eq!(tri_bvh, tri_ref, "different triangle hit");
-            assert!((t_bvh - t_ref).abs() < 1e-4, "t mismatch: {t_bvh} vs {t_ref}");
-            if tri_bvh.is_some() {
-                hits += 1;
-            }
-        }
-        assert!(hits > 20, "test rays barely hit the scene ({hits}/200)");
-    }
-
-    #[test]
-    fn test_bvh_leaf_ranges_cover_all_triangles() {
-        let mut tris = random_scene(300, 9);
-        let nodes = build_bvh(&mut tris);
-        let mut seen = vec![false; tris.len()];
-        for node in &nodes {
-            if node.count > 0 {
-                for i in node.left_first..node.left_first + node.count {
-                    assert!(!seen[i as usize], "triangle {i} in two leaves");
-                    seen[i as usize] = true;
-                }
-            }
-        }
-        assert!(seen.iter().all(|&s| s), "not every triangle is in a leaf");
-    }
-
-    #[test]
-    fn test_bvh_degenerate_identical_centroids() {
-        // All triangles share one centroid: SAH can't split, the median
-        // fallback must still terminate and cover everything.
-        let tri = RtTriangle {
-            p0: [0.0, 0.0, 0.0],
-            p1: [1.0, 0.0, 0.0],
-            p2: [0.0, 1.0, 0.0],
-            material: 0,
-        };
-        let mut tris = vec![tri; 100];
-        let nodes = build_bvh(&mut tris);
-        let covered: u32 = nodes.iter().filter(|n| n.count > 0).map(|n| n.count).sum();
-        assert_eq!(covered, 100);
-        let (t, hit) = traverse_bvh_cpu(&nodes, &tris, [0.2, 0.2, -5.0], [0.0, 0.0, 1.0]);
-        assert!(hit.is_some());
-        assert!((t - 5.0).abs() < 1e-3);
-    }
-
-    #[test]
-    fn test_bvh_empty_and_single() {
-        let mut empty: Vec<RtTriangle> = Vec::new();
-        assert!(build_bvh(&mut empty).is_empty());
-
-        let mut single = vec![RtTriangle {
-            p0: [-1.0, -1.0, 0.0],
-            p1: [1.0, -1.0, 0.0],
-            p2: [0.0, 1.0, 0.0],
-            material: 3,
-        }];
-        let nodes = build_bvh(&mut single);
-        assert_eq!(nodes.len(), 1);
-        assert_eq!(nodes[0].count, 1);
-        let (t, hit) = traverse_bvh_cpu(&nodes, &single, [0.0, 0.0, -3.0], [0.0, 0.0, 1.0]);
-        assert_eq!(hit, Some(0));
-        assert!((t - 3.0).abs() < 1e-4);
-    }
-
     #[test]
     fn test_rt_shaders_compile() {
         // naga parse + validate + SPIR-V write for both tiers; panics on failure.
-        let tier1 = compile_wgsl(&format!(
-            "{}\n{}",
-            include_str!("rt_common.wgsl"),
-            include_str!("rt_bvh.wgsl")
-        ));
+        let tier1 = compile_wgsl(&crate::draw::shaders::rt_bvh_source());
         assert!(!tier1.is_empty());
-        let tier2 = compile_wgsl_ray_query(&format!(
-            "{}\n{}",
-            include_str!("rt_common.wgsl"),
-            include_str!("rt_query.wgsl")
-        ));
+        let tier2 = compile_wgsl_ray_query(&format!("{}\n{}", crate::draw::shaders::RT_COMMON, crate::draw::shaders::RT_QUERY));
         assert!(!tier2.is_empty());
-        let denoise = compile_wgsl(include_str!("rt_denoise.wgsl"));
+        let denoise = compile_wgsl(crate::draw::shaders::RT_DENOISE);
         assert!(!denoise.is_empty());
     }
 
diff --git a/src/web/mod.rs b/src/web/mod.rs
index 97041ac..d6871da 100644
--- a/src/web/mod.rs
+++ b/src/web/mod.rs
@@ -11,6 +11,7 @@
 
 mod compute;
 mod renderer;
+mod rt;
 mod scene;
 mod shell;
 
diff --git a/src/web/renderer.rs b/src/web/renderer.rs
index c4df5c6..7ebae23 100644
--- a/src/web/renderer.rs
+++ b/src/web/renderer.rs
@@ -26,7 +26,9 @@
 
 use std::collections::HashMap;
 
+use super::rt::WebRt;
 use super::scene::WebScene;
+use crate::draw::rt::{RtCamera, RtEnvironment, RtImage, RtMaterial, RtTriangle};
 use crate::draw::scene::{MeshId, SceneDraw, SceneImage, Stage3D, Vertex3D};
 
 use wasm_bindgen::{JsCast, JsValue};
@@ -171,6 +173,11 @@ pub struct WebRenderer {
     /// the backdrop) — what the 2D pass binds while a scene is shown.
     scene: WebScene,
     scene_group_0: Option<GpuBindGroup>,
+    /// The path tracer, made by the first `set_rt_scene`; the background and
+    /// environment it is handed at each `stage_rt`, as on Vulkan.
+    rt: Option<WebRt>,
+    rt_background: Option<[f32; 3]>,
+    rt_environment: RtEnvironment,
     blocks: Growable,
     group_1: GpuBindGroup,
     vertices: Growable,
@@ -444,6 +451,9 @@ impl WebRenderer {
             snapshot: None,
             scene,
             scene_group_0: None,
+            rt: None,
+            rt_background: None,
+            rt_environment: RtEnvironment::default(),
             blocks,
             group_1,
             vertices,
@@ -610,7 +620,8 @@ impl WebRenderer {
         let (w, h) = (target.width(), target.height());
         // The scene's backdrop follows the canvas; a resized one holds no
         // scene until the next is drawn into it, as on Vulkan.
-        if self.scene.has_staged() || self.scene.target.is_some() {
+        let rt_staged = self.rt.as_ref().is_some_and(|rt| rt.staged());
+        if self.scene.has_staged() || rt_staged || self.scene.target.is_some() {
             let had = self.scene.target.as_ref().map(|t| (t.width, t.height));
             self.scene.fit(&self.device, w, h)?;
             if had != Some((w, h)) {
@@ -689,6 +700,16 @@ impl WebRenderer {
         let encoder = self.device.create_command_encoder();
         let images_for_scene = &self.images;
         self.scene.record(&self.device, &self.queue, &encoder, &|id| images_for_scene.get(&id).map(|i| i.group.clone()))?;
+        // The traced pane, into the same backdrop after the raster scene.
+        if let (Some(rt), Some(t)) = (self.rt.as_mut(), self.scene.target.as_ref()) {
+            let image_view = |id: u32| {
+                let img = images_for_scene.get(&id)?;
+                Some((img.texture.create_view().ok()?, img.width, img.height))
+            };
+            if rt.record(&self.device, &self.queue, &encoder, &t.view, (t.width, t.height), &image_view)? {
+                self.scene.backdrop_valid = true;
+            }
+        }
         let base_group_0 = match (&self.scene_group_0, &self.scene.target) {
             (Some(group), Some(t)) if self.scene.backdrop_valid => {
                 encoder.copy_texture_to_texture_with_gpu_extent_3d_dict(
@@ -861,4 +882,37 @@ impl Stage3D for WebRenderer {
             self.scene.light = v.normalize().to_array();
         }
     }
+    fn set_rt_scene_with_image(&mut self, triangles: &[RtTriangle], materials: &[RtMaterial], image: Option<RtImage>) {
+        if self.rt.is_none() {
+            match WebRt::new(&self.device, self.view_format) {
+                Ok(rt) => self.rt = Some(rt),
+                Err(e) => {
+                    web_sys::console::error_2(&"cce-ui: no path tracer:".into(), &e);
+                    return;
+                }
+            }
+        }
+        let rt = self.rt.as_mut().unwrap();
+        if let Err(e) = rt.set_scene(&self.device, &self.queue, triangles, materials, image) {
+            web_sys::console::error_2(&"cce-ui: the traced scene was not uploaded:".into(), &e);
+        }
+    }
+    fn set_rt_environment(&mut self, environment: RtEnvironment) {
+        self.rt_environment = environment;
+    }
+    fn set_rt_background(&mut self, color: Option<[f32; 3]>) {
+        self.rt_background = color;
+    }
+    fn stage_rt(&mut self, pane: (u32, u32, u32, u32), camera: RtCamera) {
+        if let Some(rt) = self.rt.as_mut() {
+            rt.set_background(self.rt_background);
+            rt.set_environment(self.rt_environment);
+            if let Err(e) = rt.stage(&self.device, pane, camera) {
+                web_sys::console::error_2(&"cce-ui: the traced pane was not staged:".into(), &e);
+            }
+        }
+    }
+    fn rt_accumulating(&self) -> bool {
+        self.rt.as_ref().is_some_and(|rt| rt.accumulating())
+    }
 }
diff --git a/src/web/rt.rs b/src/web/rt.rs
new file mode 100644
index 0000000..e0bf597
--- /dev/null
+++ b/src/web/rt.rs
@@ -0,0 +1,469 @@
+//! The path tracer on WebGPU: the Vulkan `RtStage`'s compute tier, ported.
+//! The same shaders (`draw::shaders::rt_bvh_source`, `RT_DENOISE`), the
+//! same scene packing, BVH and parameter blocks (`draw::rt`), the same
+//! rules for when the accumulation restarts — a camera, pane size, scene,
+//! background or environment change — and the same frame: one sample a
+//! frame added into the running mean, three à-trous denoise iterations over
+//! it, the result put into the backdrop's pane region in place of the
+//! raster scene, a pane that moved clearing the backdrop once first.
+//!
+//! What WebGPU makes different:
+//!
+//! - **The compute tier only.** The hardware ray-query tier needs
+//!   `VK_KHR_ray_query`; WebGPU has no ray tracing. A BVH traversed in
+//!   compute is the tier every Vulkan device without RT cores runs too.
+//! - **The blit is a draw.** Vulkan blits the tracer's `rgba8unorm` image
+//!   into the sRGB backdrop, converting as it copies; WebGPU copies only
+//!   between formats that differ in sRGB-ness at most, so a small render
+//!   pass loads each texel and writes it through the backdrop's sRGB view —
+//!   the same conversion (the unorm value taken as linear and encoded).
+//! - **No frames in flight to juggle.** One parameter buffer and one bind
+//!   group per dispatch, written before the frame's single submission.
+
+use wasm_bindgen::JsValue;
+use web_sys::{
+    gpu_buffer_usage as buffer_usage, gpu_shader_stage as shader_stage, gpu_texture_usage as texture_usage, GpuBindGroup,
+    GpuBindGroupDescriptor, GpuBindGroupEntry, GpuBindGroupLayout, GpuBindGroupLayoutDescriptor, GpuBindGroupLayoutEntry,
+    GpuBuffer, GpuBufferBinding, GpuBufferBindingLayout, GpuBufferBindingType, GpuColorTargetState, GpuCommandEncoder,
+    GpuComputePassDescriptor, GpuComputePipeline, GpuComputePipelineDescriptor, GpuDevice, GpuFragmentState, GpuLoadOp,
+    GpuPipelineLayoutDescriptor, GpuProgrammableStage, GpuQueue, GpuRenderPassColorAttachment, GpuRenderPassDescriptor,
+    GpuRenderPipeline, GpuRenderPipelineDescriptor, GpuSampler, GpuSamplerBindingLayout, GpuSamplerBindingType,
+    GpuStorageTextureAccess, GpuStorageTextureBindingLayout, GpuStoreOp, GpuTexture, GpuTextureBindingLayout,
+    GpuTextureFormat, GpuTextureSampleType, GpuTextureView, GpuTextureViewDimension, GpuVertexState,
+};
+
+use super::renderer::{shader_module, texture, whole_view};
+use crate::draw::rt::{
+    denoise_params, pack_scene, rt_params, DenoiseParams, ParamImage, RtCamera, RtEnvironment, RtImage, RtMaterial,
+    RtParams, RtTriangle, DENOISE_ITERATIONS, MAX_SAMPLES, WORKGROUP,
+};
+use crate::draw::shaders::{rt_bvh_source, RT_DENOISE};
+
+/// Each denoise iteration's block sits at a multiple of this.
+const STRIDE: u32 = 256;
+
+/// Loads the tracer's image at the fragment's place in the pane and writes
+/// it through the backdrop's sRGB view: Vulkan's UNORM-to-sRGB blit.
+const BLIT: &str = "
+struct Blit { origin: vec2<i32>, _pad: vec2<i32> }
+@group(0) @binding(0) var src: texture_2d<f32>;
+@group(0) @binding(1) var<uniform> blit: Blit;
+@vertex fn vs_main(@builtin(vertex_index) i: u32) -> @builtin(position) vec4<f32> {
+    let p = vec2<f32>(f32((i << 1u) & 2u), f32(i & 2u));
+    return vec4<f32>(p * 2.0 - 1.0, 0.0, 1.0);
+}
+@fragment fn fs_main(@builtin(position) pos: vec4<f32>) -> @location(0) vec4<f32> {
+    return textureLoad(src, vec2<i32>(pos.xy) - blit.origin, 0);
+}";
+
+/// The pane-sized targets: the running sum, the denoiser's features and
+/// ping-pong, and the image the result is written to.
+struct Targets {
+    accum: GpuBuffer,
+    features: GpuBuffer,
+    ping: GpuBuffer,
+    pong: GpuBuffer,
+    output: GpuTexture,
+    output_view: GpuTextureView,
+    /// Denoise iterations 0 and 2 (src pong, dst ping), and 1 (src ping, dst pong).
+    denoise_b: GpuBindGroup,
+    denoise_a: GpuBindGroup,
+    blit: GpuBindGroup,
+    size: (u32, u32),
+}
+
+/// The scene the buffers hold, and the image standing in it.
+struct Scene {
+    nodes: GpuBuffer,
+    tris: GpuBuffer,
+    materials: GpuBuffer,
+    tri_count: u32,
+    image: Option<RtImage>,
+}
+
+pub(crate) struct WebRt {
+    trace: GpuComputePipeline,
+    trace_layout: GpuBindGroupLayout,
+    denoise: GpuComputePipeline,
+    denoise_layout: GpuBindGroupLayout,
+    blit: GpuRenderPipeline,
+    blit_layout: GpuBindGroupLayout,
+    params: GpuBuffer,
+    denoise_params: GpuBuffer,
+    blit_params: GpuBuffer,
+    stand_in: GpuTextureView,
+    sampler: GpuSampler,
+    scene: Option<Scene>,
+    targets: Option<Targets>,
+
+    pane: (u32, u32, u32, u32),
+    pane_moved: bool,
+    camera: Option<RtCamera>,
+    sample_index: u32,
+    spp: u32,
+    staged: bool,
+    background: Option<[f32; 3]>,
+    environment: RtEnvironment,
+}
+
+fn buffer_entry(binding: u32, ty: GpuBufferBindingType, dynamic: bool) -> GpuBindGroupLayoutEntry {
+    let entry = GpuBindGroupLayoutEntry::new(binding, shader_stage::COMPUTE);
+    let layout = GpuBufferBindingLayout::new();
+    layout.set_type(ty);
+    layout.set_has_dynamic_offset(dynamic);
+    entry.set_buffer(&layout);
+    entry
+}
+
+fn storage_texture_entry(binding: u32) -> GpuBindGroupLayoutEntry {
+    let entry = GpuBindGroupLayoutEntry::new(binding, shader_stage::COMPUTE);
+    let layout = GpuStorageTextureBindingLayout::new(GpuTextureFormat::Rgba8unorm);
+    layout.set_access(GpuStorageTextureAccess::WriteOnly);
+    layout.set_view_dimension(GpuTextureViewDimension::N2d);
+    entry.set_storage_texture(&layout);
+    entry
+}
+
+fn texture_entry(binding: u32, visibility: u32, sample: GpuTextureSampleType) -> GpuBindGroupLayoutEntry {
+    let entry = GpuBindGroupLayoutEntry::new(binding, visibility);
+    let layout = GpuTextureBindingLayout::new();
+    layout.set_sample_type(sample);
+    layout.set_view_dimension(GpuTextureViewDimension::N2d);
+    entry.set_texture(&layout);
+    entry
+}
+
+fn whole(binding: u32, buffer: &GpuBuffer) -> GpuBindGroupEntry {
+    GpuBindGroupEntry::new_with_gpu_buffer_binding(binding, &GpuBufferBinding::new(buffer))
+}
+
+fn sized(binding: u32, buffer: &GpuBuffer, size: u32) -> GpuBindGroupEntry {
+    let b = GpuBufferBinding::new(buffer);
+    b.set_size(size);
+    GpuBindGroupEntry::new_with_gpu_buffer_binding(binding, &b)
+}
+
+fn buffer(device: &GpuDevice, size: u32, usage: u32, label: &str) -> Result<GpuBuffer, JsValue> {
+    let desc = web_sys::GpuBufferDescriptor::new(size.max(16).next_multiple_of(4), usage);
+    desc.set_label(label);
+    device.create_buffer(&desc)
+}
+
+fn compute_pipeline(device: &GpuDevice, layout: &GpuBindGroupLayout, code: &str, entry: &str, label: &str) -> GpuComputePipeline {
+    let pipeline_layout = device.create_pipeline_layout(&GpuPipelineLayoutDescriptor::new(&[js_sys::JsOption::wrap(layout.clone())]));
+    let stage = GpuProgrammableStage::new(&shader_module(device, code, label));
+    stage.set_entry_point(entry);
+    let desc = GpuComputePipelineDescriptor::new(&pipeline_layout, &stage);
+    desc.set_label(label);
+    device.create_compute_pipeline(&desc)
+}
+
+impl WebRt {
+    /// The tracer's pipelines; `backdrop_format` is what the blit writes.
+    pub(crate) fn new(device: &GpuDevice, backdrop_format: GpuTextureFormat) -> Result<Self, JsValue> {
+        use GpuBufferBindingType::{ReadOnlyStorage, Storage, Uniform};
+        let sampler_entry = {
+            let entry = GpuBindGroupLayoutEntry::new(8, shader_stage::COMPUTE);
+            let layout = GpuSamplerBindingLayout::new();
+            layout.set_type(GpuSamplerBindingType::Filtering);
+            entry.set_sampler(&layout);
+            entry
+        };
+        let trace_layout = device.create_bind_group_layout(&GpuBindGroupLayoutDescriptor::new(&[
+            buffer_entry(0, Uniform, false),
+            buffer_entry(1, ReadOnlyStorage, false),
+            buffer_entry(2, ReadOnlyStorage, false),
+            buffer_entry(3, ReadOnlyStorage, false),
+            buffer_entry(4, Storage, false),
+            storage_texture_entry(5),
+            buffer_entry(6, Storage, false),
+            texture_entry(7, shader_stage::COMPUTE, GpuTextureSampleType::Float),
+            sampler_entry,
+        ]))?;
+        let trace = compute_pipeline(device, &trace_layout, &rt_bvh_source(), "cs_main", "rt-trace");
+        let denoise_layout = device.create_bind_group_layout(&GpuBindGroupLayoutDescriptor::new(&[
+            buffer_entry(0, Uniform, true),
+            buffer_entry(1, ReadOnlyStorage, false),
+            buffer_entry(2, ReadOnlyStorage, false),
+            buffer_entry(3, ReadOnlyStorage, false),
+            buffer_entry(4, Storage, false),
+            storage_texture_entry(5),
+        ]))?;
+        let denoise = compute_pipeline(device, &denoise_layout, RT_DENOISE, "cs_denoise", "rt-denoise");
+
+        let blit_layout = device.create_bind_group_layout(&GpuBindGroupLayoutDescriptor::new(&[
+            texture_entry(0, shader_stage::FRAGMENT, GpuTextureSampleType::UnfilterableFloat),
+            {
+                let entry = GpuBindGroupLayoutEntry::new(1, shader_stage::FRAGMENT);
+                let layout = GpuBufferBindingLayout::new();
+                layout.set_type(GpuBufferBindingType::Uniform);
+                entry.set_buffer(&layout);
+                entry
+            },
+        ]))?;
+        let blit_module = shader_module(device, BLIT, "rt-blit");
+        let vertex = GpuVertexState::new(&blit_module);
+        vertex.set_entry_point("vs_main");
+        let targets = [js_sys::JsOption::wrap(GpuColorTargetState::new(backdrop_format))];
+        let fragment = GpuFragmentState::new(&blit_module, &targets);
+        fragment.set_entry_point("fs_main");
+        let desc = GpuRenderPipelineDescriptor::new(
+            &device.create_pipeline_layout(&GpuPipelineLayoutDescriptor::new(&[js_sys::JsOption::wrap(blit_layout.clone())])),
+            &vertex,
+        );
+        desc.set_fragment(&fragment);
+        desc.set_label("rt-blit");
+        let blit = device.create_render_pipeline(&desc)?;
+
+        let params = buffer(device, std::mem::size_of::<RtParams>() as u32, buffer_usage::UNIFORM | buffer_usage::COPY_DST, "rt-params")?;
+        let denoise_params =
+            buffer(device, STRIDE * DENOISE_ITERATIONS as u32, buffer_usage::UNIFORM | buffer_usage::COPY_DST, "rt-denoise-params")?;
+        let blit_params = buffer(device, 16, buffer_usage::UNIFORM | buffer_usage::COPY_DST, "rt-blit-params")?;
+        // The image binding's stand-in while the scene has none: one clear texel.
+        let stand_in = texture(device, GpuTextureFormat::Rgba8unormSrgb, 1, 1, texture_usage::TEXTURE_BINDING | texture_usage::COPY_DST, "rt-stand-in")?;
+        let sampler = {
+            let desc = web_sys::GpuSamplerDescriptor::new();
+            desc.set_mag_filter(web_sys::GpuFilterMode::Linear);
+            desc.set_min_filter(web_sys::GpuFilterMode::Linear);
+            desc.set_mipmap_filter(web_sys::GpuMipmapFilterMode::Linear);
+            device.create_sampler_with_descriptor(&desc)
+        };
+        Ok(Self {
+            trace,
+            trace_layout,
+            denoise,
+            denoise_layout,
+            blit,
+            blit_layout,
+            params,
+            denoise_params,
+            blit_params,
+            stand_in: whole_view(&stand_in)?,
+            sampler,
+            scene: None,
+            targets: None,
+            pane: (0, 0, 0, 0),
+            pane_moved: false,
+            camera: None,
+            sample_index: 0,
+            spp: 1,
+            staged: false,
+            background: None,
+            environment: RtEnvironment::default(),
+        })
+    }
+
+    /// Replace the scene: packed and its BVH built on the CPU, uploaded.
+    pub(crate) fn set_scene(
+        &mut self,
+        device: &GpuDevice,
+        queue: &GpuQueue,
+        triangles: &[RtTriangle],
+        materials: &[RtMaterial],
+        image: Option<RtImage>,
+    ) -> Result<(), JsValue> {
+        let packed = pack_scene(triangles, materials, image.map(|i| i.corners), true);
+        let upload = |bytes: &[u8], label: &str| -> Result<GpuBuffer, JsValue> {
+            let b = buffer(device, bytes.len() as u32, buffer_usage::STORAGE | buffer_usage::COPY_DST, label)?;
+            if !bytes.is_empty() {
+                queue.write_buffer_with_u32_and_u8_slice(&b, 0, bytes)?;
+            }
+            Ok(b)
+        };
+        if let Some(old) = self.scene.take() {
+            for b in [old.nodes, old.tris, old.materials] {
+                b.destroy();
+            }
+        }
+        self.scene = Some(Scene {
+            nodes: upload(bytemuck::cast_slice(&packed.nodes), "rt-nodes")?,
+            tris: upload(bytemuck::cast_slice(&packed.tris), "rt-tris")?,
+            materials: upload(bytemuck::cast_slice(&packed.materials), "rt-materials")?,
+            tri_count: packed.tris.len() as u32,
+            image,
+        });
+        self.sample_index = 0;
+        Ok(())
+    }
+
+    pub(crate) fn set_background(&mut self, background: Option<[f32; 3]>) {
+        if self.background != background {
+            self.background = background;
+            self.sample_index = 0;
+        }
+    }
+
+    pub(crate) fn set_environment(&mut self, environment: RtEnvironment) {
+        if self.environment != environment {
+            self.environment = environment;
+            self.sample_index = 0;
+        }
+    }
+
+    /// Stage a frame for `pane` (physical px): the targets follow its size,
+    /// and a camera or size change restarts the accumulation; a pane that
+    /// moved clears the backdrop once before the next blit.
+    pub(crate) fn stage(&mut self, device: &GpuDevice, pane: (u32, u32, u32, u32), camera: RtCamera) -> Result<(), JsValue> {
+        let (_, _, w, h) = pane;
+        if w == 0 || h == 0 {
+            self.staged = false;
+            return Ok(());
+        }
+        if self.targets.as_ref().map(|t| t.size) != Some((w, h)) {
+            self.recreate_targets(device, w, h)?;
+            self.sample_index = 0;
+        }
+        if pane != self.pane && self.pane != (0, 0, 0, 0) {
+            self.pane_moved = true;
+        }
+        if self.camera != Some(camera) {
+            self.camera = Some(camera);
+            self.sample_index = 0;
+        }
+        self.pane = pane;
+        self.staged = true;
+        Ok(())
+    }
+
+    pub(crate) fn staged(&self) -> bool {
+        self.staged
+    }
+
+    /// True while another dispatch would still refine the image.
+    pub(crate) fn accumulating(&self) -> bool {
+        self.scene.as_ref().is_some_and(|s| s.tri_count > 0) && self.sample_index < MAX_SAMPLES
+    }
+
+    fn recreate_targets(&mut self, device: &GpuDevice, w: u32, h: u32) -> Result<(), JsValue> {
+        if let Some(old) = self.targets.take() {
+            for b in [old.accum, old.features, old.ping, old.pong] {
+                b.destroy();
+            }
+            old.output.destroy();
+        }
+        let px = w * h;
+        let storage = buffer_usage::STORAGE;
+        let accum = buffer(device, px * 16, storage, "rt-accum")?;
+        let features = buffer(device, px * 32, storage, "rt-features")?;
+        let ping = buffer(device, px * 16, storage, "rt-denoise-ping")?;
+        let pong = buffer(device, px * 16, storage, "rt-denoise-pong")?;
+        let output = texture(
+            device,
+            GpuTextureFormat::Rgba8unorm,
+            w,
+            h,
+            texture_usage::STORAGE_BINDING | texture_usage::TEXTURE_BINDING,
+            "rt-output",
+        )?;
+        let output_view = whole_view(&output)?;
+        let denoise_group = |src: &GpuBuffer, dst: &GpuBuffer| {
+            let entries = [
+                sized(0, &self.denoise_params, std::mem::size_of::<DenoiseParams>() as u32),
+                whole(1, &accum),
+                whole(2, &features),
+                whole(3, src),
+                whole(4, dst),
+                GpuBindGroupEntry::new_with_gpu_texture_view(5, &output_view),
+            ];
+            device.create_bind_group(&GpuBindGroupDescriptor::new(&entries, &self.denoise_layout))
+        };
+        let denoise_b = denoise_group(&pong, &ping);
+        let denoise_a = denoise_group(&ping, &pong);
+        let blit = device.create_bind_group(&GpuBindGroupDescriptor::new(
+            &[GpuBindGroupEntry::new_with_gpu_texture_view(0, &output_view), whole(1, &self.blit_params)],
+            &self.blit_layout,
+        ));
+        self.targets = Some(Targets { accum, features, ping, pong, output, output_view, denoise_b, denoise_a, blit, size: (w, h) });
+        Ok(())
+    }
+
+    /// Record one accumulation dispatch, the denoise and the blit into
+    /// `backdrop` (the scene's, `backdrop_size` px). False when there is
+    /// nothing to do (not staged, an empty scene, converged): the backdrop
+    /// then keeps what it has. `image_view` looks the scene's image up in
+    /// the 2D pass's table: its view and size, if it has landed.
+    pub(crate) fn record(
+        &mut self,
+        device: &GpuDevice,
+        queue: &GpuQueue,
+        encoder: &GpuCommandEncoder,
+        backdrop: &GpuTextureView,
+        backdrop_size: (u32, u32),
+        image_view: &dyn Fn(u32) -> Option<(GpuTextureView, u32, u32)>,
+    ) -> Result<bool, JsValue> {
+        let ready = self.staged && self.scene.as_ref().is_some_and(|s| s.tri_count > 0) && self.targets.is_some();
+        self.staged = false;
+        if !ready || self.sample_index >= MAX_SAMPLES {
+            return Ok(false);
+        }
+        let (Some(scene), Some(t), Some(camera)) = (&self.scene, &self.targets, self.camera) else { return Ok(false) };
+        let (w, h) = t.size;
+
+        // The parameter blocks: the scene's image as it is NOW (its upload
+        // may land after the scene was set), or none, which lets rays through.
+        let bound = scene.image.and_then(|i| image_view(i.image).map(|(view, iw, ih)| (view, i, iw, ih)));
+        let (view, param_image) = match bound {
+            Some((view, i, iw, ih)) => (view, Some(ParamImage { width: iw, height: ih, corners: i.corners, opacity: i.opacity })),
+            None => (self.stand_in.clone(), None),
+        };
+        let params = rt_params(camera, (w, h), self.sample_index, self.spp, param_image, self.background, &self.environment);
+        queue.write_buffer_with_u32_and_u8_slice(&self.params, 0, bytemuck::bytes_of(&params))?;
+        let mut blocks = vec![0u8; STRIDE as usize * DENOISE_ITERATIONS];
+        for (i, p) in denoise_params(w, h, self.sample_index + self.spp).iter().enumerate() {
+            let at = i * STRIDE as usize;
+            blocks[at..at + std::mem::size_of::<DenoiseParams>()].copy_from_slice(bytemuck::bytes_of(p));
+        }
+        queue.write_buffer_with_u32_and_u8_slice(&self.denoise_params, 0, &blocks)?;
+        let (px, py, _, _) = self.pane;
+        let dst_x = px.min(backdrop_size.0);
+        let dst_y = py.min(backdrop_size.1);
+        let (bw, bh) = (w.min(backdrop_size.0 - dst_x), h.min(backdrop_size.1 - dst_y));
+        queue.write_buffer_with_u32_and_u8_slice(&self.blit_params, 0, bytemuck::cast_slice(&[dst_x as i32, dst_y as i32, 0, 0]))?;
+
+        let trace_group = device.create_bind_group(&GpuBindGroupDescriptor::new(
+            &[
+                whole(0, &self.params),
+                whole(1, &scene.nodes),
+                whole(2, &scene.tris),
+                whole(3, &scene.materials),
+                whole(4, &t.accum),
+                GpuBindGroupEntry::new_with_gpu_texture_view(5, &t.output_view),
+                whole(6, &t.features),
+                GpuBindGroupEntry::new_with_gpu_texture_view(7, &view),
+                GpuBindGroupEntry::new(8, &self.sampler),
+            ],
+            &self.trace_layout,
+        ));
+        let groups = (w.div_ceil(WORKGROUP), h.div_ceil(WORKGROUP));
+        let pass = encoder.begin_compute_pass_with_descriptor(&GpuComputePassDescriptor::new());
+        pass.set_pipeline(&self.trace);
+        pass.set_bind_group(0, Some(&trace_group));
+        pass.dispatch_workgroups_with_workgroup_count_y(groups.0, groups.1);
+        // À-trous iterations: 0 and 2 write ping, 1 writes pong; the last
+        // rewrites the output image.
+        pass.set_pipeline(&self.denoise);
+        for i in 0..DENOISE_ITERATIONS {
+            let group = if i % 2 == 0 { &t.denoise_b } else { &t.denoise_a };
+            pass.set_bind_group_with_u32_slice_and_u32_and_dynamic_offsets_data_length(0, Some(group), &[i as u32 * STRIDE], 0, 1)?;
+            pass.dispatch_workgroups_with_workgroup_count_y(groups.0, groups.1);
+        }
+        pass.end();
+
+        // Into the backdrop's pane region (clearing the whole backdrop first
+        // when the pane moved: stale pixels sit outside the new region).
+        let load = if std::mem::take(&mut self.pane_moved) { GpuLoadOp::Clear } else { GpuLoadOp::Load };
+        let color = GpuRenderPassColorAttachment::new_with_gpu_texture_view(load, GpuStoreOp::Store, backdrop);
+        color.set_clear_value(&[0.0, 0.0, 0.0, 0.0].map(js_sys::Number::from));
+        let pass = encoder.begin_render_pass(&GpuRenderPassDescriptor::new(&[js_sys::JsOption::wrap(color)]))?;
+        if bw > 0 && bh > 0 {
+            pass.set_viewport(dst_x as f32, dst_y as f32, bw as f32, bh as f32, 0.0, 1.0);
+            pass.set_scissor_rect(dst_x, dst_y, bw, bh);
+            pass.set_pipeline(&self.blit);
+            pass.set_bind_group(0, Some(&t.blit));
+            pass.draw(3);
+        }
+        pass.end();
+        self.sample_index += self.spp;
+        Ok(true)
+    }
+}