GPU-accelerated UI toolkit (Vulkan)
git clone https://git.lucas.co/cce-ui.git
W4c: the path tracer runs in the browser
Stage3D carries the tracer's half too (set_rt_scene[_with_image],
set_rt_environment, set_rt_background, stage_rt, rt_accumulating), and
what it traces from moves to draw::rt: the schema, the binned-SAH BVH
and its tests, scene packing (pack_scene) and the parameter blocks
(rt_params, denoise_params) as the shaders read them, the constants;
the rt_*.wgsl shaders move to draw/. vk::rt re-exports the schema and
keeps its own state (frames in flight, the ray-query tier,
RtOffscreen), now packing through draw::rt — its GPU tests pass on
lavapipe.
web/rt.rs is the compute tier on WebGPU: the same shaders, packing and
restart rules, a sample a frame and three a-trous iterations in one
compute pass. WebGPU has no ray tracing, so there is no ray-query tier;
and where Vulkan blits the tracer's rgba8unorm image into the sRGB
backdrop, converting as it copies, a small render pass loads each texel
and writes it through the backdrop's sRGB view.
The 3D probe gains a traced mode that stages exactly eight frames:
Vulkan's compute tier on lavapipe against SwiftShader, the traced
pane's mean differs by 0.10 of a level and every pixel is within 8.
Native is unchanged: the same two known test failures (534 pass with
the GPU tests), the plate golden byte-identical, the 24-step run
identical.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WjL3pejMNY95NHv9BcmXaZ
CLAUDE.md | 27 +-
Cargo.toml | 2 +
examples/probe3d/scene.rs | 67 +++-
examples/probe3d_native.rs | 8 +-
examples/probe3d_web.rs | 16 +-
scripts/web-probe/capture-app.mjs | 11 +-
scripts/web-probe/demo.html | 7 +-
scripts/web-probe/probe3d | 17 +-
src/draw/mod.rs | 1 +
src/draw/rt.rs | 691 ++++++++++++++++++++++++++++++++++++++
src/{vk => draw}/rt_bvh.wgsl | 0
src/{vk => draw}/rt_common.wgsl | 0
src/{vk => draw}/rt_denoise.wgsl | 0
src/{vk => draw}/rt_query.wgsl | 0
src/draw/scene.rs | 28 +-
src/draw/shaders.rs | 13 +
src/engine.rs | 1 +
src/vk/renderer.rs | 15 +
src/vk/rt.rs | 686 ++-----------------------------------
src/web/mod.rs | 1 +
src/web/renderer.rs | 56 ++-
src/web/rt.rs | 469 ++++++++++++++++++++++++++
22 files changed, 1437 insertions(+), 679 deletions(-)
diff --git a/CLAUDE.md b/CLAUDE.md
index ff7c48c..81a71bc 100644
--- a/CLAUDE.md
+++ b/CLAUDE.md
@@ -162,6 +162,28 @@ image in the scene before the translucent draw, a host light, frost over the pan
195 px differ by more than 8 levels, all on 1 px wires (where along its length a line
steps a row is the rasterizer's), everything else within 2.
+**And so does the path tracer** (since 2026-10-05). `Stage3D` carries the tracer's half
+too — `set_rt_scene` / `set_rt_scene_with_image`, `set_rt_environment`,
+`set_rt_background`, `stage_rt`, `rt_accumulating` — and what it traces from moved to
+`draw::rt`: the schema (`RtTriangle`, `RtMaterial`, `RtImage`, `RtCamera`,
+`RtEnvironment`), the binned-SAH BVH and its tests, the buffers (`pack_scene`) and
+parameter blocks (`rt_params`, `denoise_params`) as the shaders read them, and the
+constants; the shaders (`rt_common` / `rt_bvh` / `rt_query` / `rt_denoise`) moved to `draw/`.
+`vk::rt` re-exports the schema and keeps its own state (frames in flight, the ray-query
+tier, `RtOffscreen`) — it now packs and lays out through `draw::rt`, verified by its
+GPU tests (`cargo test --lib rt -- --ignored`, lavapipe). `web/rt.rs` is the compute tier on
+WebGPU: the same shaders and packing, the same rules for restarting the accumulation, one
+sample a frame plus the three à-trous iterations in one compute pass. Two differences:
+WebGPU has no ray tracing, so there is no ray-query tier (the compute tier is what every
+Vulkan device without RT cores runs too); and Vulkan BLITS the tracer's `rgba8unorm`
+image into the sRGB backdrop, converting as it copies, which WebGPU's copies cannot — a
+small render pass loads each texel and writes it through the backdrop's sRGB view, the
+same conversion. The probe's traced mode (`Probe3d<true>`, `PROBE3D_TRACE=1` natively,
+`scripts/web-probe/probe3d <out> traced`) stages exactly eight frames and stops, so both
+are compared at eight samples: 2026-10-05, Vulkan compute tier (`CCE_VK_RT=compute`) on
+lavapipe vs SwiftShader, the traced pane's mean differs by 0.10 of a level, every pixel
+within 8, 47 channels in the frame past 8 — a few paths that diverged.
+
**The reference app runs on both, through one input script.** `examples/demo_web.rs` is
`src/main.rs`'s `DemoApp` (included by `#[path]`, hence `pub(crate)`) in a page;
`scripts/web-probe/demo <dir>` builds it, serves it with the machine's fonts and replays the
@@ -1270,7 +1292,9 @@ cce-system-interface) to confirm behavior, not just the test suite.
at a per-batch dynamic offset (`WEBGPU_BLOCK_STRIDE`); the backdrop is sampled with
`textureSampleLevel(…, 0.0)` because WebGPU rejects implicit-LOD sampling in the
non-uniform blur branch (the backdrop has one level, so the texel is the same —
- `frost_pair` is identical to the pixel either way).
+ `frost_pair` is identical to the pixel either way). And the 3D halves: `draw::scene`
+ (the raster scene's types, uniforms and the `Stage3D` trait) and `draw::rt` (the path
+ tracer's schema, BVH and parameter blocks), with their shaders beside the 2D ones.
- `web/` — wasm32 only: `WebRenderer` (`new(canvas).await`, `resize`, `prepare_text`,
`draw_frame_2d`, and `capture_next_frame` / `take_capture().await` or
`take_pending_capture` to read a frame back). Its module doc lists what differs from the
@@ -1278,6 +1302,7 @@ cce-system-interface) to confirm behavior, not just the test suite.
dynamic-offset uniform, a 1x1 backdrop, the blur snapshot as end-pass / copy / resume,
every frame drawn whole. And `shell.rs`, the browser shell: `run`, `Fonts`, `Sizing`,
`capture` (see "And an `Application` runs in a page" above); `scene.rs`, the 3D pass;
+ `rt.rs`, the path tracer's compute tier;
`compute.rs`, the async
`ComputeDevice`; and `request_device`, the adapter and device every one of them asks
for (with the limits a caller names raised to the adapter's).
diff --git a/Cargo.toml b/Cargo.toml
index 3086395..962a607 100644
--- a/Cargo.toml
+++ b/Cargo.toml
@@ -179,6 +179,8 @@ web-sys = { version = "0.3", features = [
"GpuCullMode",
"GpuFrontFace",
"GpuRenderPassDepthStencilAttachment",
+ "GpuStorageTextureAccess",
+ "GpuStorageTextureBindingLayout",
] }
# The renderer probe: one scene through both renderers (examples/probe/).
diff --git a/examples/probe3d/scene.rs b/examples/probe3d/scene.rs
index 028c116..0f460d8 100644
--- a/examples/probe3d/scene.rs
+++ b/examples/probe3d/scene.rs
@@ -12,7 +12,8 @@
//! blur samples the backdrop the scene left.
use cce_ui::engine::{
- AppSender, Application, LogicalPosition, LogicalSize, MeshId, SceneDraw, SceneImage, Stage3D, Vertex3D, WindowSettings,
+ AppSender, Application, LogicalPosition, LogicalSize, MeshId, RtCamera, RtImage, RtMaterial, RtTriangle, SceneDraw,
+ SceneImage, Stage3D, Vertex3D, WindowSettings,
};
use cce_ui::scene::layout::Rect;
use cce_ui::scene::paint::{DisplayList, PaintCtx, PlateSpec};
@@ -25,11 +26,22 @@ pub const H: u32 = 800;
/// The UI strip's width; the 3D pane is the rest of the window.
const STRIP: f32 = 300.0;
-pub struct Probe3d {
+/// The traced probe's sample count: frames staged before it stops.
+pub const TRACE_FRAMES: u32 = 8;
+
+pub struct Probe3d<const TRACE: bool> {
image: u32,
meshes: Option<Meshes>,
+ /// Traced frames staged so far.
+ traced: u32,
}
+pub type Raster = Probe3d<false>;
+pub type Traced = Probe3d<true>;
+
+/// The image's corners in the scene, as both passes place it.
+const IMAGE_CORNERS: [[f32; 3]; 4] = [[-2.6, 1.6, -1.6], [-0.6, 1.6, -1.6], [-0.6, 0.35, -1.6], [-2.6, 0.35, -1.6]];
+
struct Meshes {
background: MeshId,
cube: MeshId,
@@ -112,7 +124,25 @@ fn sphere(center: Vec3, radius: f32, stacks: u32, slices: u32) -> (Vec<Vertex3D>
(tris, lines)
}
-impl Probe3d {
+/// The traced scene: the raster meshes' triangles, a material each.
+fn traced_scene() -> (Vec<RtTriangle>, Vec<RtMaterial>) {
+ let mut tris = Vec::new();
+ let mut mats = Vec::new();
+ let mut add = |verts: Vec<Vertex3D>, albedo: [f32; 3], emission: [f32; 3]| {
+ let m = mats.len() as u32;
+ mats.push(RtMaterial { albedo, emission });
+ for t in verts.chunks_exact(3) {
+ tris.push(RtTriangle { p0: t[0].position, p1: t[1].position, p2: t[2].position, material: m });
+ }
+ };
+ add(cuboid(Vec3::new(0.0, -0.75, 0.0), Vec3::new(3.0, 0.05, 2.4), [[0.3; 3]; 6]), [0.55, 0.57, 0.55], [0.0; 3]);
+ add(cuboid(Vec3::new(-1.6, 0.0, 0.2), Vec3::splat(0.6), [[0.0; 3]; 6]), [0.8, 0.35, 0.25], [0.0; 3]);
+ add(sphere(Vec3::new(0.4, 0.3, -0.4), 0.9, 12, 20).0, [0.45, 0.55, 0.75], [0.0; 3]);
+ add(cuboid(Vec3::new(1.5, 0.1, 1.2), Vec3::splat(0.55), [[0.0; 3]; 6]), [0.3, 0.7, 0.9], [0.4, 0.9, 1.1]);
+ (tris, mats)
+}
+
+impl<const TRACE: bool> Probe3d<TRACE> {
fn camera(size: LogicalSize, scale: f64) -> (Mat4, (u32, u32, u32, u32)) {
let s = scale as f32;
let (pw, ph) = ((size.width - STRIP) * s, size.height * s);
@@ -124,9 +154,19 @@ impl Probe3d {
let shift = Mat4::from_translation(Vec3::new(1.0 - ndc_w, 0.0, 0.0)) * Mat4::from_scale(Vec3::new(ndc_w, 1.0, 1.0));
(shift * proj * view, ((STRIP * s) as u32, 0, pw as u32, ph as u32))
}
+
+ /// The traced pane's camera: the pane's own projection, unshifted —
+ /// the tracer's image is the pane.
+ fn trace_camera(size: LogicalSize, scale: f64) -> RtCamera {
+ let s = scale as f32;
+ let (pw, ph) = ((size.width - STRIP) * s, size.height * s);
+ let proj = Mat4::perspective_rh(40f32.to_radians(), pw / ph, 0.1, 100.0);
+ let view = Mat4::look_at_rh(Vec3::new(3.2, 2.4, 5.6), Vec3::new(0.0, 0.2, 0.0), Vec3::Y);
+ RtCamera { inv_mvp: (proj * view).inverse().to_cols_array_2d() }
+ }
}
-impl Application for Probe3d {
+impl<const TRACE: bool> Application for Probe3d<TRACE> {
type Message = ();
fn create(_sender: AppSender<()>) -> Self {
@@ -138,7 +178,7 @@ impl Application for Probe3d {
px.extend([(x * 255 / (iw - 1)) as u8, (y * 255 / (ih - 1)) as u8, 200, a]);
}
}
- Probe3d { image: cce_ui::draw::upload_rgba(px, iw, ih), meshes: None }
+ Probe3d { image: cce_ui::draw::upload_rgba(px, iw, ih), meshes: None, traced: 0 }
}
fn settings(&self) -> WindowSettings {
@@ -181,9 +221,24 @@ impl Application for Probe3d {
let glass_wires = stage.create_mesh(&cuboid_edges(glass_c, Vec3::splat(0.55), [0.9, 0.95, 1.0]));
stage.set_scene_light([0.6, 0.7, 0.4]);
self.meshes = Some(Meshes { background, cube, prelit, sphere, sphere_wires, glass, glass_wires });
+ if TRACE {
+ let (tris, mats) = traced_scene();
+ stage.set_rt_scene_with_image(&tris, &mats, Some(RtImage { image: self.image, corners: IMAGE_CORNERS, opacity: 0.9 }));
+ }
}
fn stage_3d(&mut self, stage: &mut dyn Stage3D, size: LogicalSize, scale: f64) -> bool {
+ if TRACE {
+ // A sample a frame for TRACE_FRAMES frames, then the backdrop
+ // keeps the result.
+ if self.traced >= TRACE_FRAMES {
+ return false;
+ }
+ let (_, pane) = Self::camera(size, scale);
+ stage.stage_rt(pane, Self::trace_camera(size, scale));
+ self.traced += 1;
+ return self.traced < TRACE_FRAMES;
+ }
let Some(m) = &self.meshes else { return false };
let (mvp, scissor) = Self::camera(size, scale);
let mvp = mvp.to_cols_array_2d();
@@ -210,7 +265,7 @@ impl Application for Probe3d {
stage.stage_scene(scissor, draws);
stage.stage_scene_images(vec![SceneImage {
image: self.image,
- corners: [[-2.6, 1.6, -1.6], [-0.6, 1.6, -1.6], [-0.6, 0.35, -1.6], [-2.6, 0.35, -1.6]],
+ corners: IMAGE_CORNERS,
mvp,
opacity: 0.9,
before: 5,
diff --git a/examples/probe3d_native.rs b/examples/probe3d_native.rs
index 5d3935f..fd91f84 100644
--- a/examples/probe3d_native.rs
+++ b/examples/probe3d_native.rs
@@ -1,9 +1,15 @@
//! The 3D probe (`probe3d/scene.rs`) on the Vulkan renderer, in the Wayland
//! runner: screenshot it and compare it with `probe3d_web`'s frame.
+//! `PROBE3D_TRACE=1` runs the traced probe (`CCE_VK_RT=compute` keeps
+//! Vulkan on the compute tier the browser has).
#[path = "probe3d/scene.rs"]
mod scene;
fn main() {
- cce_ui::engine::run::<scene::Probe3d>();
+ if std::env::var("PROBE3D_TRACE").is_ok_and(|v| v == "1") {
+ cce_ui::engine::run::<scene::Traced>();
+ } else {
+ cce_ui::engine::run::<scene::Raster>();
+ }
}
diff --git a/examples/probe3d_web.rs b/examples/probe3d_web.rs
index 139e41e..a868f52 100644
--- a/examples/probe3d_web.rs
+++ b/examples/probe3d_web.rs
@@ -12,11 +12,21 @@ mod web {
use cce_ui::web::{Fonts, Sizing};
use wasm_bindgen::prelude::*;
+ fn fonts(fonts: js_sys::Array) -> Fonts {
+ std::panic::set_hook(Box::new(|info| web_sys::console::error_1(&info.to_string().into())));
+ Fonts::new(fonts.iter().map(|f| js_sys::Uint8Array::new(&f).to_vec()).collect())
+ }
+
+ /// The raster probe.
#[wasm_bindgen]
pub async fn start(canvas: web_sys::HtmlCanvasElement, fonts: js_sys::Array, _families: String, _fill: bool) -> Result<(), JsValue> {
- std::panic::set_hook(Box::new(|info| web_sys::console::error_1(&info.to_string().into())));
- let fonts = Fonts::new(fonts.iter().map(|f| js_sys::Uint8Array::new(&f).to_vec()).collect());
- cce_ui::web::run::<super::scene::Probe3d>(canvas, fonts, Sizing::Page).await
+ cce_ui::web::run::<super::scene::Raster>(canvas, self::fonts(fonts), Sizing::Page).await
+ }
+
+ /// The traced probe (`demo.html?app=probe3d_web&entry=start_traced`).
+ #[wasm_bindgen]
+ pub async fn start_traced(canvas: web_sys::HtmlCanvasElement, fonts: js_sys::Array, _families: String, _fill: bool) -> Result<(), JsValue> {
+ cce_ui::web::run::<super::scene::Traced>(canvas, self::fonts(fonts), Sizing::Page).await
}
#[wasm_bindgen]
diff --git a/scripts/web-probe/capture-app.mjs b/scripts/web-probe/capture-app.mjs
index 7b62ee6..ecff41a 100644
--- a/scripts/web-probe/capture-app.mjs
+++ b/scripts/web-probe/capture-app.mjs
@@ -1,13 +1,14 @@
-// capture-app.mjs <site-dir> <app> <out.rgba> [settle-ms]: open demo.html
-// with ?app=<app> (browser.mjs), let the app run for settle-ms (default
-// 3000), and write the frame `capture()` reads back from the GPU to
+// capture-app.mjs <site-dir> <app> <out.rgba> [settle-ms] [entry]: open
+// demo.html with ?app=<app> (and &entry=<entry>, the app's start function in
+// place of `start`) through browser.mjs, let the app run for settle-ms
+// (default 3000), and write the frame `capture()` reads back from the GPU to
// <out.rgba> (premultiplied RGBA8, 1280x800).
import fs from 'fs';
import { open } from './browser.mjs';
-const [root, app, out, settle] = process.argv.slice(2);
+const [root, app, out, settle, entry] = process.argv.slice(2);
if (!root || !app || !out) { console.error('usage: capture-app.mjs <site-dir> <app> <out.rgba> [settle-ms]'); process.exit(2); }
-const { page, logs, close } = await open(root, 'demo.html', 1280, 800, `?app=${app}`);
+const { page, logs, close } = await open(root, 'demo.html', 1280, 800, `?app=${app}` + (entry ? `&entry=${entry}` : ''));
let ok = true;
try {
await page.waitForFunction(() => window.demoReady === true, null, { timeout: 120000 });
diff --git a/scripts/web-probe/demo.html b/scripts/web-probe/demo.html
index a01c1b9..f1813c7 100644
--- a/scripts/web-probe/demo.html
+++ b/scripts/web-probe/demo.html
@@ -3,7 +3,8 @@
<title>cce-ui</title>
<!-- An app on the browser shell: examples/demo_web.rs, or the example the
query names (?app=probe3d_web) — any that exports start(canvas, fonts,
- families, fill) and capture(). The page owns the canvas's size (fill =
+ families, fill), or the entry the query names in its place
+ (&entry=start_traced), and capture(). The page owns the canvas's size (fill =
Sizing::Page). window.capture() reads the next frame back from the GPU,
base64 RGBA8: a headless Chromium leaves a WebGPU canvas out of its
screenshots. -->
@@ -11,7 +12,9 @@
<canvas id="c"></canvas>
<script type="module">
const app = new URLSearchParams(location.search).get('app') || 'demo_web';
-const { default: init, start, capture } = await import(`./${app}.js`);
+const mod = await import(`./${app}.js`);
+const { default: init, capture } = mod;
+const start = mod[new URLSearchParams(location.search).get('entry') || 'start'];
await init();
const { files, families } = await (await fetch('fonts.json')).json();
const fonts = [];
diff --git a/scripts/web-probe/probe3d b/scripts/web-probe/probe3d
index e117efe..a1494ef 100755
--- a/scripts/web-probe/probe3d
+++ b/scripts/web-probe/probe3d
@@ -1,7 +1,9 @@
#!/usr/bin/env bash
-# web-probe/probe3d <out.rgba> — run the 3D probe (examples/probe3d/scene.rs)
-# on the browser shell and WebGPU renderer in headless Chromium, and write a
-# frame read back from the GPU: 1280x800 premultiplied RGBA8.
+# web-probe/probe3d <out.rgba> [traced] — run the 3D probe
+# (examples/probe3d/scene.rs) on the browser shell and WebGPU renderer in
+# headless Chromium, and write a frame read back from the GPU: 1280x800
+# premultiplied RGBA8. With `traced`, the path-traced probe, given time to
+# take its eight samples ($PROBE3D_SETTLE ms, default 60000).
#
# Its other half is `cargo run --example probe3d_native`, screenshotted in a
# compositor; compare.py composites this frame over that compositor's
@@ -10,7 +12,8 @@
# `run` needs.
set -euo pipefail
here=$(cd "$(dirname "$0")" && pwd)
-out=$(realpath -m "${1:?usage: web-probe/probe3d <out.rgba>}")
+out=$(realpath -m "${1:?usage: web-probe/probe3d <out.rgba> [traced]}")
+mode=${2:-}
fonts=${PROBE_FONTS_DIR:-/usr/share/fonts/truetype/dejavu}
cd "$here/../.."
cargo build --release --target wasm32-unknown-unknown --example probe3d_web
@@ -22,5 +25,9 @@ cp "$here/demo.html" "$site/demo.html"
mkdir "$site/fonts"
cp "$fonts"/*.ttf "$site/fonts/"
(cd "$site/fonts" && printf '%s\n' *.ttf | python3 -c 'import json,sys; print(json.dumps({"files": sys.stdin.read().split(), "families": ""}))') > "$site/fonts.json"
-node "$here/capture-app.mjs" "$site" probe3d_web "$out"
+if [ "$mode" = traced ]; then
+ node "$here/capture-app.mjs" "$site" probe3d_web "$out" "${PROBE3D_SETTLE:-60000}" start_traced
+else
+ node "$here/capture-app.mjs" "$site" probe3d_web "$out"
+fi
echo "web-probe: wrote $out"
diff --git a/src/draw/mod.rs b/src/draw/mod.rs
index 50270ce..b096ce1 100644
--- a/src/draw/mod.rs
+++ b/src/draw/mod.rs
@@ -14,6 +14,7 @@ use cosmic_text::Buffer as TextBuffer;
pub mod glyphs;
pub mod images;
+pub mod rt;
pub mod scene;
pub mod shaders;
diff --git a/src/draw/rt.rs b/src/draw/rt.rs
new file mode 100644
index 0000000..4845067
--- /dev/null
+++ b/src/draw/rt.rs
@@ -0,0 +1,691 @@
+//! The path tracer's scene and the work done on the CPU for it, with no GPU
+//! in it: the schema an app fills ([`RtTriangle`], [`RtMaterial`],
+//! [`RtImage`], [`RtCamera`], [`RtEnvironment`]), the binned-SAH BVH built
+//! over it, the buffers and parameter blocks laid out as `rt_common.wgsl` /
+//! `rt_bvh.wgsl` / `rt_denoise.wgsl` read them, and the tracer's constants.
+//! The Vulkan stage (`vk::rt`, with a hardware ray-query tier besides) and
+//! the WebGPU one (`web::rt`, the compute tier) both trace from these.
+//!
+//! The tracer is plain compute: a CPU-built BVH traversed per pixel,
+//! progressive accumulation of one sample a frame, an à-trous denoise over
+//! the running mean, blitted into the backdrop's pane region in place of the
+//! raster scene. Scene schema is internal by design — importers (OBJ/glTF)
+//! belong in a loader that converts *into* [`RtTriangle`]/[`RtMaterial`].
+
+/// One triangle of an RT scene, in the same space as the camera's `inv_mvp`
+/// (for the designer: mesh space, the space `Vertex3D` positions live in).
+#[derive(Debug, Clone, Copy, PartialEq)]
+pub struct RtTriangle {
+ pub p0: [f32; 3],
+ pub p1: [f32; 3],
+ pub p2: [f32; 3],
+ /// Index into the material slice passed alongside.
+ pub material: u32,
+}
+
+/// Lambertian surface + optional emission, linear color (matching the raster
+/// path, whose vertex colors land in the sRGB attachment as linear values).
+#[derive(Debug, Clone, Copy, PartialEq)]
+pub struct RtMaterial {
+ pub albedo: [f32; 3],
+ pub emission: [f32; 3],
+}
+
+/// An image standing in the traced scene: a quad that shows the image's
+/// colour as it is — unlit, as the raster pass's `SceneImage` draws it — and
+/// lets a ray through where the image is clear. A path ends on it, so to
+/// the rest of the scene it is a light of its own colour. The tracer adds
+/// the quad to the scene itself.
+///
+/// The image is one the 2D pass already holds (an id from `upload_rgba`),
+/// so a picture shown in the raster viewport costs the tracer nothing more.
+#[derive(Debug, Clone, Copy, PartialEq)]
+pub struct RtImage {
+ pub image: u32,
+ /// The quad's corners in the scene's space, in the image's own order:
+ /// top-left, top-right, bottom-right, bottom-left.
+ pub corners: [[f32; 3]; 4],
+ /// Alpha multiplier over the image's own.
+ pub opacity: f32,
+}
+
+/// [`RtImage`] for the headless tracer, which has no 2D pass to share an
+/// image with and is handed the pixels: tightly packed sRGB RGBA8.
+#[derive(Debug, Clone, Copy)]
+pub struct RtImagePixels<'a> {
+ pub pixels: &'a [u8],
+ pub width: u32,
+ pub height: u32,
+ pub corners: [[f32; 3]; 4],
+ pub opacity: f32,
+}
+
+/// The full camera: the inverse of the raster path's `proj * view * model`.
+/// Rays are unprojected from NDC through it, so any matrix stack that renders
+/// the raster viewport drives the tracer unchanged.
+#[derive(Debug, Clone, Copy, PartialEq)]
+pub struct RtCamera {
+ pub inv_mvp: [[f32; 4]; 4],
+}
+
+// --- GPU layouts (must match rt.wgsl) ---
+
+#[repr(C)]
+#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
+pub(crate) struct GpuTriangle {
+ pub(crate) p0: [f32; 4], // w = material index (bitcast)
+ pub(crate) p1: [f32; 4],
+ pub(crate) p2: [f32; 4],
+}
+
+#[repr(C)]
+#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
+pub(crate) struct GpuMaterial {
+ pub(crate) albedo: [f32; 4],
+ pub(crate) emission: [f32; 4],
+}
+
+#[repr(C)]
+#[derive(Debug, Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
+pub(crate) struct GpuBvhNode {
+ pub(crate) min: [f32; 3],
+ /// Leaf (`count > 0`): first triangle. Internal: left child; right = +1.
+ pub(crate) left_first: u32,
+ pub(crate) max: [f32; 3],
+ pub(crate) count: u32,
+}
+
+#[repr(C)]
+#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
+pub(crate) struct RtParams {
+ pub(crate) inv_mvp: [[f32; 4]; 4],
+ pub(crate) width: u32,
+ pub(crate) height: u32,
+ pub(crate) sample_index: u32,
+ pub(crate) max_bounces: u32,
+ pub(crate) spp: u32,
+ pub(crate) _pad: [u32; 3],
+ pub(crate) img_origin: [f32; 4],
+ pub(crate) img_u: [f32; 4],
+ pub(crate) img_v: [f32; 4],
+ pub(crate) background: [f32; 4],
+ // The environment, each xyz in a vec4 (w unused): toward the sun
+ // (unit), the sun's radiance, the sky overhead, the sky below.
+ pub(crate) sun_dir: [f32; 4],
+ pub(crate) sun_color: [f32; 4],
+ pub(crate) sky_zenith: [f32; 4],
+ pub(crate) sky_nadir: [f32; 4],
+}
+
+/// The traced scene's light: a sky that grades from `sky_nadir` straight
+/// down to `sky_zenith` straight up, and a sun — a bright lobe toward
+/// `sun_direction` of `sun_color` radiance. It is the tracer's ONLY light
+/// (nothing in a scene emits unless its material does). Colours are linear
+/// RGB and may exceed 1. [`Default`] is the studio sky the tracer always
+/// had.
+#[derive(Clone, Copy, Debug, PartialEq)]
+pub struct RtEnvironment {
+ /// Toward the sun, world space; any length (zero is straight up).
+ pub sun_direction: [f32; 3],
+ pub sun_color: [f32; 3],
+ pub sky_zenith: [f32; 3],
+ pub sky_nadir: [f32; 3],
+}
+
+impl Default for RtEnvironment {
+ fn default() -> Self {
+ Self {
+ sun_direction: [0.45, 0.75, 0.35],
+ sun_color: [8.0, 7.6, 6.8],
+ sky_zenith: [0.72, 0.82, 0.98],
+ sky_nadir: [0.32, 0.31, 0.35],
+ }
+ }
+}
+
+impl RtEnvironment {
+ pub(crate) fn sun_unit(&self) -> [f32; 3] {
+ let v = glam::Vec3::from_array(self.sun_direction);
+ if v.length_squared() > 1e-12 { v.normalize().to_array() } else { [0.0, 1.0, 0.0] }
+ }
+}
+
+// --- BVH construction (binned SAH) ---
+
+const BVH_BINS: usize = 8;
+const BVH_LEAF_MAX: u32 = 4;
+
+#[derive(Clone, Copy)]
+struct Aabb {
+ min: [f32; 3],
+ max: [f32; 3],
+}
+
+impl Aabb {
+ const EMPTY: Aabb = Aabb { min: [f32::INFINITY; 3], max: [f32::NEG_INFINITY; 3] };
+
+ fn grow(&mut self, p: [f32; 3]) {
+ for a in 0..3 {
+ self.min[a] = self.min[a].min(p[a]);
+ self.max[a] = self.max[a].max(p[a]);
+ }
+ }
+
+ fn grow_aabb(&mut self, other: &Aabb) {
+ self.grow(other.min);
+ self.grow(other.max);
+ }
+
+ fn half_area(&self) -> f32 {
+ let dx = (self.max[0] - self.min[0]).max(0.0);
+ let dy = (self.max[1] - self.min[1]).max(0.0);
+ let dz = (self.max[2] - self.min[2]).max(0.0);
+ dx * dy + dy * dz + dz * dx
+ }
+}
+
+fn tri_aabb(t: &RtTriangle) -> Aabb {
+ let mut b = Aabb::EMPTY;
+ b.grow(t.p0);
+ b.grow(t.p1);
+ b.grow(t.p2);
+ b
+}
+
+fn tri_centroid(t: &RtTriangle) -> [f32; 3] {
+ let mut c = [0.0f32; 3];
+ for a in 0..3 {
+ c[a] = (t.p0[a] + t.p1[a] + t.p2[a]) / 3.0;
+ }
+ c
+}
+
+/// Build a BVH over `triangles`, reordering them so leaves reference
+/// contiguous ranges. Returns the flat node array (empty input → empty vec).
+pub(crate) fn build_bvh(triangles: &mut Vec<RtTriangle>) -> Vec<GpuBvhNode> {
+ if triangles.is_empty() {
+ return Vec::new();
+ }
+ let bounds: Vec<Aabb> = triangles.iter().map(tri_aabb).collect();
+ let centroids: Vec<[f32; 3]> = triangles.iter().map(tri_centroid).collect();
+ let mut order: Vec<u32> = (0..triangles.len() as u32).collect();
+
+ fn range_bounds(order: &[u32], bounds: &[Aabb]) -> Aabb {
+ let mut b = Aabb::EMPTY;
+ for &i in order {
+ b.grow_aabb(&bounds[i as usize]);
+ }
+ b
+ }
+
+ let mut nodes: Vec<GpuBvhNode> = Vec::with_capacity(triangles.len() * 2);
+ let root_bounds = range_bounds(&order, &bounds);
+ nodes.push(GpuBvhNode {
+ min: root_bounds.min,
+ left_first: 0,
+ max: root_bounds.max,
+ count: triangles.len() as u32,
+ });
+
+ // (node index, start, count) work list over `order`.
+ let mut work = vec![(0usize, 0usize, triangles.len())];
+ while let Some((node_idx, start, count)) = work.pop() {
+ if (count as u32) <= BVH_LEAF_MAX {
+ continue; // stays a leaf
+ }
+ let slice = &mut order[start..start + count];
+
+ // Centroid bounds pick the split axis.
+ let mut cb = Aabb::EMPTY;
+ for &i in slice.iter() {
+ cb.grow(centroids[i as usize]);
+ }
+ let mut axis = 0;
+ let mut extent = 0.0f32;
+ for a in 0..3 {
+ let e = cb.max[a] - cb.min[a];
+ if e > extent {
+ extent = e;
+ axis = a;
+ }
+ }
+
+ let mut split_at = None;
+ if extent > 1e-12 {
+ // Binned SAH along `axis`.
+ let scale = BVH_BINS as f32 / extent;
+ let bin_of = |i: u32| -> usize {
+ (((centroids[i as usize][axis] - cb.min[axis]) * scale) as usize)
+ .min(BVH_BINS - 1)
+ };
+ let mut bin_bounds = [Aabb::EMPTY; BVH_BINS];
+ let mut bin_counts = [0usize; BVH_BINS];
+ for &i in slice.iter() {
+ let b = bin_of(i);
+ bin_counts[b] += 1;
+ bin_bounds[b].grow_aabb(&bounds[i as usize]);
+ }
+ // Cost of each of the BINS-1 split planes.
+ let mut best_cost = f32::INFINITY;
+ let mut best_plane = 0usize;
+ for plane in 1..BVH_BINS {
+ let (mut lb, mut rb) = (Aabb::EMPTY, Aabb::EMPTY);
+ let (mut lc, mut rc) = (0usize, 0usize);
+ for b in 0..plane {
+ lb.grow_aabb(&bin_bounds[b]);
+ lc += bin_counts[b];
+ }
+ for b in plane..BVH_BINS {
+ rb.grow_aabb(&bin_bounds[b]);
+ rc += bin_counts[b];
+ }
+ if lc == 0 || rc == 0 {
+ continue;
+ }
+ let cost = lb.half_area() * lc as f32 + rb.half_area() * rc as f32;
+ if cost < best_cost {
+ best_cost = cost;
+ best_plane = plane;
+ }
+ }
+ if best_plane > 0 {
+ let mut mid = 0usize;
+ for k in 0..count {
+ if bin_of(slice[k]) < best_plane {
+ slice.swap(k, mid);
+ mid += 1;
+ }
+ }
+ if mid > 0 && mid < count {
+ split_at = Some(mid);
+ }
+ }
+ }
+ // Degenerate centroids or a one-sided SAH result: median split keeps
+ // the tree balanced instead of forcing a giant leaf.
+ let mid = split_at.unwrap_or(count / 2);
+
+ let left_bounds = range_bounds(&slice[..mid], &bounds);
+ let right_bounds = range_bounds(&slice[mid..], &bounds);
+ let left_idx = nodes.len();
+ nodes.push(GpuBvhNode {
+ min: left_bounds.min,
+ left_first: (start) as u32,
+ max: left_bounds.max,
+ count: mid as u32,
+ });
+ nodes.push(GpuBvhNode {
+ min: right_bounds.min,
+ left_first: (start + mid) as u32,
+ max: right_bounds.max,
+ count: (count - mid) as u32,
+ });
+ nodes[node_idx].left_first = left_idx as u32;
+ nodes[node_idx].count = 0;
+ work.push((left_idx, start, mid));
+ work.push((left_idx + 1, start + mid, count - mid));
+ }
+
+ // Apply the final order to the triangle array so leaf ranges are direct.
+ let reordered: Vec<RtTriangle> =
+ order.iter().map(|&i| triangles[i as usize]).collect();
+ *triangles = reordered;
+ nodes
+}
+
+// --- The Vulkan stage ---
+
+pub(crate) const MAX_SAMPLES: u32 = 1024;
+pub(crate) const MAX_BOUNCES: u32 = 4;
+pub(crate) const WORKGROUP: u32 = 8;
+
+#[repr(C)]
+#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
+pub(crate) struct DenoiseParams {
+ pub(crate) width: u32,
+ pub(crate) height: u32,
+ pub(crate) step: u32,
+ pub(crate) first: u32,
+ pub(crate) last: u32,
+ pub(crate) inv_sqrt_n: f32,
+ pub(crate) _pad: [u32; 2],
+}
+
+pub(crate) const DENOISE_ITERATIONS: usize = 3; // à-trous steps 1, 2, 4
+
+/// A scene as the tracer's buffers hold it.
+pub(crate) struct PackedScene {
+ pub tris: Vec<GpuTriangle>,
+ pub materials: Vec<GpuMaterial>,
+ /// Empty unless `with_bvh` was asked for (the ray-query tier builds
+ /// its own structure).
+ pub nodes: Vec<GpuBvhNode>,
+}
+
+/// Lay a scene out for the tracer. An image joins it as a quad of two
+/// triangles (corners in its own order, top-left first) under a material
+/// of its own, marked textured (`albedo.w`), past the scene's materials —
+/// or past the one stood in for a scene that names none — so every tier
+/// meets it as it meets any triangle, and a scene that is an image alone is
+/// not an empty one. `with_bvh` builds the BVH, reordering the triangles.
+pub(crate) fn pack_scene(
+ triangles: &[RtTriangle],
+ materials: &[RtMaterial],
+ image_corners: Option<[[f32; 3]; 4]>,
+ with_bvh: bool,
+) -> PackedScene {
+ let mut tris: Vec<RtTriangle> = triangles.to_vec();
+ let image_material = materials.len().max(1) as u32;
+ if let Some([tl, tr, br, bl]) = image_corners {
+ tris.push(RtTriangle { p0: tl, p1: bl, p2: tr, material: image_material });
+ tris.push(RtTriangle { p0: tr, p1: bl, p2: br, material: image_material });
+ }
+ let nodes = if with_bvh { build_bvh(&mut tris) } else { Vec::new() };
+ let gpu_tris = tris
+ .iter()
+ .map(|t| GpuTriangle {
+ p0: [t.p0[0], t.p0[1], t.p0[2], f32::from_bits(t.material)],
+ p1: [t.p1[0], t.p1[1], t.p1[2], 0.0],
+ p2: [t.p2[0], t.p2[1], t.p2[2], 0.0],
+ })
+ .collect();
+ let mut gpu_mats: Vec<GpuMaterial> = if materials.is_empty() {
+ vec![GpuMaterial { albedo: [0.8, 0.8, 0.8, 0.0], emission: [0.0; 4] }]
+ } else {
+ materials
+ .iter()
+ .map(|m| GpuMaterial {
+ albedo: [m.albedo[0], m.albedo[1], m.albedo[2], 0.0],
+ emission: [m.emission[0], m.emission[1], m.emission[2], 0.0],
+ })
+ .collect()
+ };
+ if image_corners.is_some() {
+ // albedo.w marks it textured: the shader takes the colour from the image.
+ gpu_mats.push(GpuMaterial { albedo: [1.0, 1.0, 1.0, 1.0], emission: [0.0; 4] });
+ }
+ PackedScene { tris: gpu_tris, materials: gpu_mats, nodes }
+}
+
+/// The image a frame traces, as its parameter block describes it: the
+/// texture's size and the quad's corners and opacity.
+pub(crate) struct ParamImage {
+ pub width: u32,
+ pub height: u32,
+ pub corners: [[f32; 3]; 4],
+ pub opacity: f32,
+}
+
+fn v4([x, y, z]: [f32; 3]) -> [f32; 4] {
+ [x, y, z, 0.0]
+}
+
+/// One dispatch's parameter block. With no image (or one not uploaded
+/// yet) the quad's opacity is 0, which lets every ray through it.
+#[allow(clippy::too_many_arguments)]
+pub(crate) fn rt_params(
+ camera: RtCamera,
+ size: (u32, u32),
+ sample_index: u32,
+ spp: u32,
+ image: Option<ParamImage>,
+ background: Option<[f32; 3]>,
+ environment: &RtEnvironment,
+) -> RtParams {
+ let (img_origin, img_u, img_v) = match image {
+ Some(ParamImage { width, height, corners: [tl, tr, _, bl], opacity }) => (
+ [tl[0], tl[1], tl[2], opacity.clamp(0.0, 1.0)],
+ [tr[0] - tl[0], tr[1] - tl[1], tr[2] - tl[2], width as f32],
+ [bl[0] - tl[0], bl[1] - tl[1], bl[2] - tl[2], height as f32],
+ ),
+ None => ([0.0; 4], [1.0, 0.0, 0.0, 1.0], [0.0, 1.0, 0.0, 1.0]),
+ };
+ RtParams {
+ inv_mvp: camera.inv_mvp,
+ width: size.0,
+ height: size.1,
+ sample_index,
+ max_bounces: MAX_BOUNCES,
+ spp,
+ _pad: [0; 3],
+ img_origin,
+ img_u,
+ img_v,
+ background: match background {
+ Some([r, g, b]) => [r, g, b, 1.0],
+ None => [0.0; 4],
+ },
+ sun_dir: v4(environment.sun_unit()),
+ sun_color: v4(environment.sun_color),
+ sky_zenith: v4(environment.sky_zenith),
+ sky_nadir: v4(environment.sky_nadir),
+ }
+}
+
+/// The denoiser's blocks, one per à-trous iteration (steps 1, 2, 4), for an
+/// accumulation that will hold `n_after` samples once this frame's dispatch
+/// lands — the colour sigma tightens as it grows.
+pub(crate) fn denoise_params(width: u32, height: u32, n_after: u32) -> [DenoiseParams; DENOISE_ITERATIONS] {
+ let inv_sqrt_n = 1.0 / (n_after.max(1) as f32).sqrt();
+ std::array::from_fn(|i| DenoiseParams {
+ width,
+ height,
+ step: 1 << i,
+ first: (i == 0) as u32,
+ last: (i == DENOISE_ITERATIONS - 1) as u32,
+ inv_sqrt_n,
+ _pad: [0; 2],
+ })
+}
+
+#[cfg(test)]
+mod tests {
+ use super::*;
+
+ // CPU mirror of the shader's traversal, for parity testing.
+ fn intersect_tri_cpu(ro: [f32; 3], rd: [f32; 3], t: &RtTriangle, t_limit: f32) -> f32 {
+ let sub = |a: [f32; 3], b: [f32; 3]| [a[0] - b[0], a[1] - b[1], a[2] - b[2]];
+ let cross = |a: [f32; 3], b: [f32; 3]| {
+ [
+ a[1] * b[2] - a[2] * b[1],
+ a[2] * b[0] - a[0] * b[2],
+ a[0] * b[1] - a[1] * b[0],
+ ]
+ };
+ let dot = |a: [f32; 3], b: [f32; 3]| a[0] * b[0] + a[1] * b[1] + a[2] * b[2];
+ let e1 = sub(t.p1, t.p0);
+ let e2 = sub(t.p2, t.p0);
+ let h = cross(rd, e2);
+ let a = dot(e1, h);
+ if a.abs() < 1e-8 {
+ return 1e30;
+ }
+ let f = 1.0 / a;
+ let s = sub(ro, t.p0);
+ let u = f * dot(s, h);
+ if !(0.0..=1.0).contains(&u) {
+ return 1e30;
+ }
+ let q = cross(s, e1);
+ let v = f * dot(rd, q);
+ if v < 0.0 || u + v > 1.0 {
+ return 1e30;
+ }
+ let tt = f * dot(e2, q);
+ if tt > 1e-4 && tt < t_limit {
+ return tt;
+ }
+ 1e30
+ }
+
+ fn traverse_bvh_cpu(
+ nodes: &[GpuBvhNode],
+ tris: &[RtTriangle],
+ ro: [f32; 3],
+ rd: [f32; 3],
+ ) -> (f32, Option<usize>) {
+ if nodes.is_empty() {
+ return (1e30, None);
+ }
+ let inv = [1.0 / rd[0], 1.0 / rd[1], 1.0 / rd[2]];
+ let hit_aabb = |min: [f32; 3], max: [f32; 3], t_limit: f32| -> bool {
+ let mut tn = f32::NEG_INFINITY;
+ let mut tf = f32::INFINITY;
+ for a in 0..3 {
+ let t1 = (min[a] - ro[a]) * inv[a];
+ let t2 = (max[a] - ro[a]) * inv[a];
+ tn = tn.max(t1.min(t2));
+ tf = tf.min(t1.max(t2));
+ }
+ tf >= tn.max(0.0) && tn < t_limit
+ };
+ let mut best = 1e30f32;
+ let mut best_tri = None;
+ let mut stack = vec![0u32];
+ while let Some(idx) = stack.pop() {
+ let node = &nodes[idx as usize];
+ if !hit_aabb(node.min, node.max, best) {
+ continue;
+ }
+ if node.count > 0 {
+ for i in node.left_first..node.left_first + node.count {
+ let t = intersect_tri_cpu(ro, rd, &tris[i as usize], best);
+ if t < best {
+ best = t;
+ best_tri = Some(i as usize);
+ }
+ }
+ } else {
+ stack.push(node.left_first);
+ stack.push(node.left_first + 1);
+ }
+ }
+ (best, best_tri)
+ }
+
+ fn brute_force(tris: &[RtTriangle], ro: [f32; 3], rd: [f32; 3]) -> (f32, Option<usize>) {
+ let mut best = 1e30f32;
+ let mut best_tri = None;
+ for (i, t) in tris.iter().enumerate() {
+ let tt = intersect_tri_cpu(ro, rd, t, best);
+ if tt < best {
+ best = tt;
+ best_tri = Some(i);
+ }
+ }
+ (best, best_tri)
+ }
+
+ // Deterministic LCG so the test needs no rand dependency.
+ struct Lcg(u64);
+ impl Lcg {
+ fn next_f32(&mut self) -> f32 {
+ self.0 = self.0.wrapping_mul(6364136223846793005).wrapping_add(1442695040888963407);
+ ((self.0 >> 33) as f32) / (u32::MAX >> 1) as f32
+ }
+ fn point(&mut self, scale: f32) -> [f32; 3] {
+ [
+ (self.next_f32() - 0.5) * scale,
+ (self.next_f32() - 0.5) * scale,
+ (self.next_f32() - 0.5) * scale,
+ ]
+ }
+ }
+
+ fn random_scene(n: usize, seed: u64) -> Vec<RtTriangle> {
+ let mut rng = Lcg(seed);
+ (0..n)
+ .map(|i| {
+ let c = rng.point(20.0);
+ let jitter = |rng: &mut Lcg, c: [f32; 3]| {
+ let d = rng.point(2.0);
+ [c[0] + d[0], c[1] + d[1], c[2] + d[2]]
+ };
+ RtTriangle {
+ p0: jitter(&mut rng, c),
+ p1: jitter(&mut rng, c),
+ p2: jitter(&mut rng, c),
+ material: (i % 5) as u32,
+ }
+ })
+ .collect()
+ }
+
+ #[test]
+ fn test_bvh_matches_brute_force() {
+ let mut tris = random_scene(500, 42);
+ let nodes = build_bvh(&mut tris);
+ assert!(!nodes.is_empty());
+ let mut rng = Lcg(7);
+ let mut hits = 0;
+ for _ in 0..200 {
+ let ro = rng.point(40.0);
+ let target = rng.point(10.0);
+ let d = [target[0] - ro[0], target[1] - ro[1], target[2] - ro[2]];
+ let len = (d[0] * d[0] + d[1] * d[1] + d[2] * d[2]).sqrt().max(1e-6);
+ let rd = [d[0] / len, d[1] / len, d[2] / len];
+ let (t_bvh, tri_bvh) = traverse_bvh_cpu(&nodes, &tris, ro, rd);
+ let (t_ref, tri_ref) = brute_force(&tris, ro, rd);
+ assert_eq!(tri_bvh, tri_ref, "different triangle hit");
+ assert!((t_bvh - t_ref).abs() < 1e-4, "t mismatch: {t_bvh} vs {t_ref}");
+ if tri_bvh.is_some() {
+ hits += 1;
+ }
+ }
+ assert!(hits > 20, "test rays barely hit the scene ({hits}/200)");
+ }
+
+ #[test]
+ fn test_bvh_leaf_ranges_cover_all_triangles() {
+ let mut tris = random_scene(300, 9);
+ let nodes = build_bvh(&mut tris);
+ let mut seen = vec![false; tris.len()];
+ for node in &nodes {
+ if node.count > 0 {
+ for i in node.left_first..node.left_first + node.count {
+ assert!(!seen[i as usize], "triangle {i} in two leaves");
+ seen[i as usize] = true;
+ }
+ }
+ }
+ assert!(seen.iter().all(|&s| s), "not every triangle is in a leaf");
+ }
+
+ #[test]
+ fn test_bvh_degenerate_identical_centroids() {
+ // All triangles share one centroid: SAH can't split, the median
+ // fallback must still terminate and cover everything.
+ let tri = RtTriangle {
+ p0: [0.0, 0.0, 0.0],
+ p1: [1.0, 0.0, 0.0],
+ p2: [0.0, 1.0, 0.0],
+ material: 0,
+ };
+ let mut tris = vec![tri; 100];
+ let nodes = build_bvh(&mut tris);
+ let covered: u32 = nodes.iter().filter(|n| n.count > 0).map(|n| n.count).sum();
+ assert_eq!(covered, 100);
+ let (t, hit) = traverse_bvh_cpu(&nodes, &tris, [0.2, 0.2, -5.0], [0.0, 0.0, 1.0]);
+ assert!(hit.is_some());
+ assert!((t - 5.0).abs() < 1e-3);
+ }
+
+ #[test]
+ fn test_bvh_empty_and_single() {
+ let mut empty: Vec<RtTriangle> = Vec::new();
+ assert!(build_bvh(&mut empty).is_empty());
+
+ let mut single = vec![RtTriangle {
+ p0: [-1.0, -1.0, 0.0],
+ p1: [1.0, -1.0, 0.0],
+ p2: [0.0, 1.0, 0.0],
+ material: 3,
+ }];
+ let nodes = build_bvh(&mut single);
+ assert_eq!(nodes.len(), 1);
+ assert_eq!(nodes[0].count, 1);
+ let (t, hit) = traverse_bvh_cpu(&nodes, &single, [0.0, 0.0, -3.0], [0.0, 0.0, 1.0]);
+ assert_eq!(hit, Some(0));
+ assert!((t - 3.0).abs() < 1e-4);
+ }
+}
diff --git a/src/vk/rt_bvh.wgsl b/src/draw/rt_bvh.wgsl
similarity index 100%
rename from src/vk/rt_bvh.wgsl
rename to src/draw/rt_bvh.wgsl
diff --git a/src/vk/rt_common.wgsl b/src/draw/rt_common.wgsl
similarity index 100%
rename from src/vk/rt_common.wgsl
rename to src/draw/rt_common.wgsl
diff --git a/src/vk/rt_denoise.wgsl b/src/draw/rt_denoise.wgsl
similarity index 100%
rename from src/vk/rt_denoise.wgsl
rename to src/draw/rt_denoise.wgsl
diff --git a/src/vk/rt_query.wgsl b/src/draw/rt_query.wgsl
similarity index 100%
rename from src/vk/rt_query.wgsl
rename to src/draw/rt_query.wgsl
diff --git a/src/draw/scene.rs b/src/draw/scene.rs
index 9124662..7e7ff1a 100644
--- a/src/draw/scene.rs
+++ b/src/draw/scene.rs
@@ -8,7 +8,9 @@
//! image, scissored to a pane, which the renderer copies beneath the UI pass
//! and the 2D shader's blur plates sample. A staged scene is drawn once; the
//! backdrop it leaves is shown under every later frame until the next one.
-//! `scene3d.wgsl` / `scene3d_image.wgsl` (in `draw/`) are the shaders.
+//! `scene3d.wgsl` / `scene3d_image.wgsl` (in `draw/`) are the shaders. A
+//! traced pane ([`super::rt`]) takes the raster scene's place in the same
+//! backdrop, staged through the same trait.
/// Layout-identical to the app's `geometry::Vertex3D` (bytemuck-castable at cutover).
#[repr(C)]
@@ -128,6 +130,8 @@ pub(crate) struct SceneUniforms {
/// it with its INWARD derivative normal, said the right way round.
pub(crate) const DEFAULT_SCENE_LIGHT: [f32; 3] = [0.55, -0.45, -0.7];
+use super::rt::{RtCamera, RtEnvironment, RtImage, RtMaterial, RtTriangle};
+
/// What an app stages a 3D scene through: the renderer's half of the
/// pass, the same on every renderer (`vk::VkRenderer`, `web::WebRenderer`).
/// An app takes one in `Application::init_3d` (make its meshes) and
@@ -146,6 +150,28 @@ pub trait Stage3D {
/// The direction TOWARD the flat shading's light, in world space; set
/// it once, and light a `prelit` mesh by the same vector.
fn set_scene_light(&mut self, toward: [f32; 3]);
+
+ /// Replace the path tracer's scene (triangles in the space the camera's
+ /// `inv_mvp` unprojects into); the BVH is built on the CPU. Rare: a
+ /// geometry rebuild. Restarts the accumulation.
+ fn set_rt_scene(&mut self, triangles: &[RtTriangle], materials: &[RtMaterial]) {
+ self.set_rt_scene_with_image(triangles, materials, None);
+ }
+ /// [`set_rt_scene`](Self::set_rt_scene) with an uploaded image standing
+ /// in the scene (the picture the raster pass draws as a `SceneImage`).
+ fn set_rt_scene_with_image(&mut self, triangles: &[RtTriangle], materials: &[RtMaterial], image: Option<RtImage>);
+ /// The traced scene's sky and sun. A change restarts the accumulation.
+ fn set_rt_environment(&mut self, environment: RtEnvironment);
+ /// What a camera ray that meets nothing shows (linear RGB), or `None`
+ /// for the sky. A change restarts the accumulation.
+ fn set_rt_background(&mut self, color: Option<[f32; 3]>);
+ /// Stage one progressive pass into `pane` (physical px) for the next
+ /// frame, in place of the raster scene there. Call it every frame while
+ /// tracing: each adds a sample; a camera, pane or scene change restarts.
+ fn stage_rt(&mut self, pane: (u32, u32, u32, u32), camera: RtCamera);
+ /// True while another staged frame would still refine the traced image
+ /// — the app's cue to keep asking for frames.
+ fn rt_accumulating(&self) -> bool;
}
/// A staged scene's uniform blocks, in the order its draws use them: one per
diff --git a/src/draw/shaders.rs b/src/draw/shaders.rs
index 4568d0c..8ad7000 100644
--- a/src/draw/shaders.rs
+++ b/src/draw/shaders.rs
@@ -20,6 +20,19 @@ pub const SCENE3D: &str = include_str!("scene3d.wgsl");
/// The 3D scene pass's images: textured quads under the meshes' uniforms.
pub const SCENE3D_IMAGE: &str = include_str!("scene3d_image.wgsl");
+/// The path tracer ([`super::rt`]): its shared core, then one trace tier
+/// after it — the compute BVH traversal, or (Vulkan only) hardware ray
+/// queries — and the à-trous denoiser that runs over its output.
+pub const RT_COMMON: &str = include_str!("rt_common.wgsl");
+pub const RT_BVH: &str = include_str!("rt_bvh.wgsl");
+pub const RT_QUERY: &str = include_str!("rt_query.wgsl");
+pub const RT_DENOISE: &str = include_str!("rt_denoise.wgsl");
+
+/// The compute tier's tracer: [`RT_COMMON`] with [`RT_BVH`] after it.
+pub fn rt_bvh_source() -> String {
+ format!("{RT_COMMON}\n{RT_BVH}")
+}
+
/// The one line that differs between the Vulkan and the WebGPU 2D shader.
const PUSH_BLOCK: &str = "var<push_constant> rrect_clip: RRectClip;";
const UNIFORM_BLOCK: &str = "@group(1) @binding(0) var<uniform> rrect_clip: RRectClip;";
diff --git a/src/engine.rs b/src/engine.rs
index 09c9092..6669e62 100644
--- a/src/engine.rs
+++ b/src/engine.rs
@@ -5,6 +5,7 @@
pub use crate::backend::app::{
Application, AppSender, LogicalPosition, LogicalSize, RenderContext, Stage3D, WindowAction, WindowSettings,
};
+pub use crate::draw::rt::{RtCamera, RtEnvironment, RtImage, RtMaterial, RtTriangle};
pub use crate::draw::scene::{MeshId, SceneDraw, SceneImage, Vertex3D};
pub use crate::backend::driver::PressedKey;
pub use crate::backend::tessellate::{
diff --git a/src/vk/renderer.rs b/src/vk/renderer.rs
index d032430..8e0768d 100644
--- a/src/vk/renderer.rs
+++ b/src/vk/renderer.rs
@@ -2346,6 +2346,21 @@ impl crate::draw::scene::Stage3D for VkRenderer {
fn set_scene_light(&mut self, toward: [f32; 3]) {
VkRenderer::set_scene_light(self, toward)
}
+ fn set_rt_scene_with_image(&mut self, triangles: &[RtTriangle], materials: &[RtMaterial], image: Option<RtImage>) {
+ VkRenderer::set_rt_scene_with_image(self, triangles, materials, image)
+ }
+ fn set_rt_environment(&mut self, environment: RtEnvironment) {
+ VkRenderer::set_rt_environment(self, environment)
+ }
+ fn set_rt_background(&mut self, color: Option<[f32; 3]>) {
+ VkRenderer::set_rt_background(self, color)
+ }
+ fn stage_rt(&mut self, pane: (u32, u32, u32, u32), camera: RtCamera) {
+ VkRenderer::stage_rt(self, pane, camera)
+ }
+ fn rt_accumulating(&self) -> bool {
+ VkRenderer::rt_accumulating(self)
+ }
}
impl Drop for VkRenderer {
diff --git a/src/vk/rt.rs b/src/vk/rt.rs
index 361162b..540c597 100644
--- a/src/vk/rt.rs
+++ b/src/vk/rt.rs
@@ -19,54 +19,11 @@ use gpu_allocator::MemoryLocation;
use super::renderer::{
compile_wgsl, compile_wgsl_ray_query, create_cpu_buffer, destroy_cpu_buffer, AllocatedBuffer,
};
-
-/// One triangle of an RT scene, in the same space as the camera's `inv_mvp`
-/// (for the designer: mesh space, the space `Vertex3D` positions live in).
-#[derive(Debug, Clone, Copy, PartialEq)]
-pub struct RtTriangle {
- pub p0: [f32; 3],
- pub p1: [f32; 3],
- pub p2: [f32; 3],
- /// Index into the material slice passed alongside.
- pub material: u32,
-}
-
-/// Lambertian surface + optional emission, linear color (matching the raster
-/// path, whose vertex colors land in the sRGB attachment as linear values).
-#[derive(Debug, Clone, Copy, PartialEq)]
-pub struct RtMaterial {
- pub albedo: [f32; 3],
- pub emission: [f32; 3],
-}
-
-/// An image standing in the traced scene: a quad that shows the image's
-/// colour as it is — unlit, as the raster pass's `SceneImage` draws it — and
-/// lets a ray through where the image is clear. A path ends on it, so to
-/// the rest of the scene it is a light of its own colour. The tracer adds
-/// the quad to the scene itself.
-///
-/// The image is one the 2D pass already holds (an id from `upload_rgba`),
-/// so a picture shown in the raster viewport costs the tracer nothing more.
-#[derive(Debug, Clone, Copy, PartialEq)]
-pub struct RtImage {
- pub image: u32,
- /// The quad's corners in the scene's space, in the image's own order:
- /// top-left, top-right, bottom-right, bottom-left.
- pub corners: [[f32; 3]; 4],
- /// Alpha multiplier over the image's own.
- pub opacity: f32,
-}
-
-/// [`RtImage`] for the headless tracer, which has no 2D pass to share an
-/// image with and is handed the pixels: tightly packed sRGB RGBA8.
-#[derive(Debug, Clone, Copy)]
-pub struct RtImagePixels<'a> {
- pub pixels: &'a [u8],
- pub width: u32,
- pub height: u32,
- pub corners: [[f32; 3]; 4],
- pub opacity: f32,
-}
+pub use crate::draw::rt::{RtCamera, RtEnvironment, RtImage, RtImagePixels, RtMaterial, RtTriangle};
+use crate::draw::rt::{
+ denoise_params, pack_scene, rt_params, DenoiseParams, ParamImage, RtParams, DENOISE_ITERATIONS, MAX_SAMPLES,
+ WORKGROUP,
+};
/// Where the stage's image comes from.
pub(crate) enum RtImageSource<'a> {
@@ -77,285 +34,6 @@ pub(crate) enum RtImageSource<'a> {
Pixels { pixels: &'a [u8], width: u32, height: u32 },
}
-/// The full camera: the inverse of the raster path's `proj * view * model`.
-/// Rays are unprojected from NDC through it, so any matrix stack that renders
-/// the raster viewport drives the tracer unchanged.
-#[derive(Debug, Clone, Copy, PartialEq)]
-pub struct RtCamera {
- pub inv_mvp: [[f32; 4]; 4],
-}
-
-// --- GPU layouts (must match rt.wgsl) ---
-
-#[repr(C)]
-#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
-struct GpuTriangle {
- p0: [f32; 4], // w = material index (bitcast)
- p1: [f32; 4],
- p2: [f32; 4],
-}
-
-#[repr(C)]
-#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
-struct GpuMaterial {
- albedo: [f32; 4],
- emission: [f32; 4],
-}
-
-#[repr(C)]
-#[derive(Debug, Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
-pub(crate) struct GpuBvhNode {
- pub(crate) min: [f32; 3],
- /// Leaf (`count > 0`): first triangle. Internal: left child; right = +1.
- pub(crate) left_first: u32,
- pub(crate) max: [f32; 3],
- pub(crate) count: u32,
-}
-
-#[repr(C)]
-#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
-struct RtParams {
- inv_mvp: [[f32; 4]; 4],
- width: u32,
- height: u32,
- sample_index: u32,
- max_bounces: u32,
- spp: u32,
- _pad: [u32; 3],
- img_origin: [f32; 4],
- img_u: [f32; 4],
- img_v: [f32; 4],
- background: [f32; 4],
- // The environment, each xyz in a vec4 (w unused): toward the sun
- // (unit), the sun's radiance, the sky overhead, the sky below.
- sun_dir: [f32; 4],
- sun_color: [f32; 4],
- sky_zenith: [f32; 4],
- sky_nadir: [f32; 4],
-}
-
-/// The traced scene's light: a sky that grades from `sky_nadir` straight
-/// down to `sky_zenith` straight up, and a sun — a bright lobe toward
-/// `sun_direction` of `sun_color` radiance. It is the tracer's ONLY light
-/// (nothing in a scene emits unless its material does). Colours are linear
-/// RGB and may exceed 1. [`Default`] is the studio sky the tracer always
-/// had.
-#[derive(Clone, Copy, Debug, PartialEq)]
-pub struct RtEnvironment {
- /// Toward the sun, world space; any length (zero is straight up).
- pub sun_direction: [f32; 3],
- pub sun_color: [f32; 3],
- pub sky_zenith: [f32; 3],
- pub sky_nadir: [f32; 3],
-}
-
-impl Default for RtEnvironment {
- fn default() -> Self {
- Self {
- sun_direction: [0.45, 0.75, 0.35],
- sun_color: [8.0, 7.6, 6.8],
- sky_zenith: [0.72, 0.82, 0.98],
- sky_nadir: [0.32, 0.31, 0.35],
- }
- }
-}
-
-impl RtEnvironment {
- fn sun_unit(&self) -> [f32; 3] {
- let v = glam::Vec3::from_array(self.sun_direction);
- if v.length_squared() > 1e-12 { v.normalize().to_array() } else { [0.0, 1.0, 0.0] }
- }
-}
-
-// --- BVH construction (binned SAH) ---
-
-const BVH_BINS: usize = 8;
-const BVH_LEAF_MAX: u32 = 4;
-
-#[derive(Clone, Copy)]
-struct Aabb {
- min: [f32; 3],
- max: [f32; 3],
-}
-
-impl Aabb {
- const EMPTY: Aabb = Aabb { min: [f32::INFINITY; 3], max: [f32::NEG_INFINITY; 3] };
-
- fn grow(&mut self, p: [f32; 3]) {
- for a in 0..3 {
- self.min[a] = self.min[a].min(p[a]);
- self.max[a] = self.max[a].max(p[a]);
- }
- }
-
- fn grow_aabb(&mut self, other: &Aabb) {
- self.grow(other.min);
- self.grow(other.max);
- }
-
- fn half_area(&self) -> f32 {
- let dx = (self.max[0] - self.min[0]).max(0.0);
- let dy = (self.max[1] - self.min[1]).max(0.0);
- let dz = (self.max[2] - self.min[2]).max(0.0);
- dx * dy + dy * dz + dz * dx
- }
-}
-
-fn tri_aabb(t: &RtTriangle) -> Aabb {
- let mut b = Aabb::EMPTY;
- b.grow(t.p0);
- b.grow(t.p1);
- b.grow(t.p2);
- b
-}
-
-fn tri_centroid(t: &RtTriangle) -> [f32; 3] {
- let mut c = [0.0f32; 3];
- for a in 0..3 {
- c[a] = (t.p0[a] + t.p1[a] + t.p2[a]) / 3.0;
- }
- c
-}
-
-/// Build a BVH over `triangles`, reordering them so leaves reference
-/// contiguous ranges. Returns the flat node array (empty input → empty vec).
-pub(crate) fn build_bvh(triangles: &mut Vec<RtTriangle>) -> Vec<GpuBvhNode> {
- if triangles.is_empty() {
- return Vec::new();
- }
- let bounds: Vec<Aabb> = triangles.iter().map(tri_aabb).collect();
- let centroids: Vec<[f32; 3]> = triangles.iter().map(tri_centroid).collect();
- let mut order: Vec<u32> = (0..triangles.len() as u32).collect();
-
- fn range_bounds(order: &[u32], bounds: &[Aabb]) -> Aabb {
- let mut b = Aabb::EMPTY;
- for &i in order {
- b.grow_aabb(&bounds[i as usize]);
- }
- b
- }
-
- let mut nodes: Vec<GpuBvhNode> = Vec::with_capacity(triangles.len() * 2);
- let root_bounds = range_bounds(&order, &bounds);
- nodes.push(GpuBvhNode {
- min: root_bounds.min,
- left_first: 0,
- max: root_bounds.max,
- count: triangles.len() as u32,
- });
-
- // (node index, start, count) work list over `order`.
- let mut work = vec![(0usize, 0usize, triangles.len())];
- while let Some((node_idx, start, count)) = work.pop() {
- if (count as u32) <= BVH_LEAF_MAX {
- continue; // stays a leaf
- }
- let slice = &mut order[start..start + count];
-
- // Centroid bounds pick the split axis.
- let mut cb = Aabb::EMPTY;
- for &i in slice.iter() {
- cb.grow(centroids[i as usize]);
- }
- let mut axis = 0;
- let mut extent = 0.0f32;
- for a in 0..3 {
- let e = cb.max[a] - cb.min[a];
- if e > extent {
- extent = e;
- axis = a;
- }
- }
-
- let mut split_at = None;
- if extent > 1e-12 {
- // Binned SAH along `axis`.
- let scale = BVH_BINS as f32 / extent;
- let bin_of = |i: u32| -> usize {
- (((centroids[i as usize][axis] - cb.min[axis]) * scale) as usize)
- .min(BVH_BINS - 1)
- };
- let mut bin_bounds = [Aabb::EMPTY; BVH_BINS];
- let mut bin_counts = [0usize; BVH_BINS];
- for &i in slice.iter() {
- let b = bin_of(i);
- bin_counts[b] += 1;
- bin_bounds[b].grow_aabb(&bounds[i as usize]);
- }
- // Cost of each of the BINS-1 split planes.
- let mut best_cost = f32::INFINITY;
- let mut best_plane = 0usize;
- for plane in 1..BVH_BINS {
- let (mut lb, mut rb) = (Aabb::EMPTY, Aabb::EMPTY);
- let (mut lc, mut rc) = (0usize, 0usize);
- for b in 0..plane {
- lb.grow_aabb(&bin_bounds[b]);
- lc += bin_counts[b];
- }
- for b in plane..BVH_BINS {
- rb.grow_aabb(&bin_bounds[b]);
- rc += bin_counts[b];
- }
- if lc == 0 || rc == 0 {
- continue;
- }
- let cost = lb.half_area() * lc as f32 + rb.half_area() * rc as f32;
- if cost < best_cost {
- best_cost = cost;
- best_plane = plane;
- }
- }
- if best_plane > 0 {
- let mut mid = 0usize;
- for k in 0..count {
- if bin_of(slice[k]) < best_plane {
- slice.swap(k, mid);
- mid += 1;
- }
- }
- if mid > 0 && mid < count {
- split_at = Some(mid);
- }
- }
- }
- // Degenerate centroids or a one-sided SAH result: median split keeps
- // the tree balanced instead of forcing a giant leaf.
- let mid = split_at.unwrap_or(count / 2);
-
- let left_bounds = range_bounds(&slice[..mid], &bounds);
- let right_bounds = range_bounds(&slice[mid..], &bounds);
- let left_idx = nodes.len();
- nodes.push(GpuBvhNode {
- min: left_bounds.min,
- left_first: (start) as u32,
- max: left_bounds.max,
- count: mid as u32,
- });
- nodes.push(GpuBvhNode {
- min: right_bounds.min,
- left_first: (start + mid) as u32,
- max: right_bounds.max,
- count: (count - mid) as u32,
- });
- nodes[node_idx].left_first = left_idx as u32;
- nodes[node_idx].count = 0;
- work.push((left_idx, start, mid));
- work.push((left_idx + 1, start + mid, count - mid));
- }
-
- // Apply the final order to the triangle array so leaf ranges are direct.
- let reordered: Vec<RtTriangle> =
- order.iter().map(|&i| triangles[i as usize]).collect();
- *triangles = reordered;
- nodes
-}
-
-// --- The Vulkan stage ---
-
-const MAX_SAMPLES: u32 = 1024;
-const MAX_BOUNCES: u32 = 4;
-const WORKGROUP: u32 = 8;
-
/// The two trace backends. They share every shader line except
/// `intersect_scene` (rt_bvh.wgsl vs rt_query.wgsl) and binding 1.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
@@ -560,20 +238,6 @@ struct StagedImage {
opacity: f32,
}
-#[repr(C)]
-#[derive(Clone, Copy, bytemuck::Pod, bytemuck::Zeroable)]
-struct DenoiseParams {
- width: u32,
- height: u32,
- step: u32,
- first: u32,
- last: u32,
- inv_sqrt_n: f32,
- _pad: [u32; 2],
-}
-
-const DENOISE_ITERATIONS: usize = 3; // à-trous steps 1, 2, 4
-
struct DenoiserFrame {
/// DENOISE_ITERATIONS dynamic-offset slices of [`DenoiseParams`].
uniforms: AllocatedBuffer,
@@ -649,7 +313,7 @@ impl Denoiser {
None,
)
.expect("Failed to create denoise pipeline layout");
- let spirv = compile_wgsl(include_str!("rt_denoise.wgsl"));
+ let spirv = compile_wgsl(crate::draw::shaders::RT_DENOISE);
let shader_module = device
.create_shader_module(&vk::ShaderModuleCreateInfo::default().code(&spirv), None)
.expect("Failed to create denoise shader module");
@@ -790,22 +454,12 @@ impl Denoiser {
/// sample count the accumulation will hold once this frame's dispatch
/// lands — the color sigma tightens as it grows.
fn write_frame_uniforms(&mut self, frame_index: usize, width: u32, height: u32, n_after: u32) {
- let inv_sqrt_n = 1.0 / (n_after.max(1) as f32).sqrt();
let frame = &mut self.frames[frame_index];
let mapped = frame.uniforms.allocation.as_mut().unwrap().mapped_slice_mut().unwrap();
- for i in 0..DENOISE_ITERATIONS {
- let params = DenoiseParams {
- width,
- height,
- step: 1 << i,
- first: (i == 0) as u32,
- last: (i == DENOISE_ITERATIONS - 1) as u32,
- inv_sqrt_n,
- _pad: [0; 2],
- };
+ for (i, params) in denoise_params(width, height, n_after).iter().enumerate() {
let offset = self.uniform_stride as usize * i;
mapped[offset..offset + std::mem::size_of::<DenoiseParams>()]
- .copy_from_slice(bytemuck::bytes_of(¶ms));
+ .copy_from_slice(bytemuck::bytes_of(params));
}
}
@@ -914,10 +568,6 @@ pub(crate) struct RtStage {
environment: RtEnvironment,
}
-fn v4([x, y, z]: [f32; 3]) -> [f32; 4] {
- [x, y, z, 0.0]
-}
-
impl RtStage {
/// `accel_loader` present means the device has the ray-query stack; the
/// stage then runs tier 2 unless `CCE_VK_RT=compute` forces the BVH tier.
@@ -1013,16 +663,8 @@ impl RtStage {
.expect("Failed to create RT pipeline layout");
let spirv = match tier {
- RtTier::Compute => compile_wgsl(&format!(
- "{}\n{}",
- include_str!("rt_common.wgsl"),
- include_str!("rt_bvh.wgsl")
- )),
- RtTier::RayQuery => compile_wgsl_ray_query(&format!(
- "{}\n{}",
- include_str!("rt_common.wgsl"),
- include_str!("rt_query.wgsl")
- )),
+ RtTier::Compute => compile_wgsl(&crate::draw::shaders::rt_bvh_source()),
+ RtTier::RayQuery => compile_wgsl_ray_query(&format!("{}\n{}", crate::draw::shaders::RT_COMMON, crate::draw::shaders::RT_QUERY)),
};
let shader_module = device
.create_shader_module(&vk::ShaderModuleCreateInfo::default().code(&spirv), None)
@@ -1249,17 +891,10 @@ impl RtStage {
materials: &[RtMaterial],
image: Option<(RtImageSource, [[f32; 3]; 4], f32)>,
) {
- let mut tris: Vec<RtTriangle> = triangles.to_vec();
- // The index the image's material will have, past the scene's own
- // (or past the one stood in for a scene that names none).
- let image_material = materials.len().max(1) as u32;
- if let Some((_, [tl, tr, br, bl], _)) = &image {
- tris.push(RtTriangle { p0: *tl, p1: *bl, p2: *tr, material: image_material });
- tris.push(RtTriangle { p0: *tr, p1: *bl, p2: *br, material: image_material });
- }
if let Some(mut old) = self.image.take().and_then(|i| i.owned) {
old.destroy(device, allocator);
}
+ let corners = image.as_ref().map(|(_, corners, _)| *corners);
self.image = image.map(|(source, corners, opacity)| match source {
RtImageSource::Shared(id) => {
StagedImage { shared: Some(id), owned: None, corners, opacity }
@@ -1279,34 +914,8 @@ impl RtStage {
opacity,
},
});
- let nodes = match self.tier {
- RtTier::Compute => build_bvh(&mut tris),
- RtTier::RayQuery => Vec::new(),
- };
- let gpu_tris: Vec<GpuTriangle> = tris
- .iter()
- .map(|t| GpuTriangle {
- p0: [t.p0[0], t.p0[1], t.p0[2], f32::from_bits(t.material)],
- p1: [t.p1[0], t.p1[1], t.p1[2], 0.0],
- p2: [t.p2[0], t.p2[1], t.p2[2], 0.0],
- })
- .collect();
- let mut gpu_mats: Vec<GpuMaterial> = if materials.is_empty() {
- vec![GpuMaterial { albedo: [0.8, 0.8, 0.8, 0.0], emission: [0.0; 4] }]
- } else {
- materials
- .iter()
- .map(|m| GpuMaterial {
- albedo: [m.albedo[0], m.albedo[1], m.albedo[2], 0.0],
- emission: [m.emission[0], m.emission[1], m.emission[2], 0.0],
- })
- .collect()
- };
- if self.image.is_some() {
- // albedo.w marks it textured: the shader takes the colour from
- // the image.
- gpu_mats.push(GpuMaterial { albedo: [1.0, 1.0, 1.0, 1.0], emission: [0.0; 4] });
- }
+ let packed = pack_scene(triangles, materials, corners, self.tier == RtTier::Compute);
+ let (gpu_tris, gpu_mats, nodes) = (packed.tris, packed.materials, packed.nodes);
self.destroy_accel(device, allocator);
for buf in [&mut self.nodes, &mut self.tris, &mut self.materials] {
@@ -1347,7 +956,7 @@ impl RtStage {
vk::BufferUsageFlags::STORAGE_BUFFER,
"rt-materials",
);
- self.tri_count = tris.len() as u32;
+ self.tri_count = gpu_tris.len() as u32;
self.sample_index = 0;
// Binding 1 (per tier), then the shared 2/3.
@@ -1884,18 +1493,13 @@ impl RtStage {
(None, Some(id)) => shared(id)?,
(None, None) => return None,
};
- Some((view, w, h, image.corners, image.opacity))
+ Some((view, ParamImage { width: w, height: h, corners: image.corners, opacity: image.opacity }))
});
- let (view, img_origin, img_u, img_v) = match bound {
- Some((view, w, h, [tl, tr, _, bl], opacity)) => (
- view,
- [tl[0], tl[1], tl[2], opacity.clamp(0.0, 1.0)],
- [tr[0] - tl[0], tr[1] - tl[1], tr[2] - tl[2], w as f32],
- [bl[0] - tl[0], bl[1] - tl[1], bl[2] - tl[2], h as f32],
- ),
- // No image, or one not there yet: opacity 0 lets every ray
- // through the quad, and the stand-in is never sampled for it.
- None => (self.stand_in.view, [0.0; 4], [1.0, 0.0, 0.0, 1.0], [0.0, 1.0, 0.0, 1.0]),
+ // No image, or one not there yet: opacity 0 lets every ray through
+ // the quad, and the stand-in is never sampled for it.
+ let (view, param_image) = match bound {
+ Some((view, p)) => (view, Some(p)),
+ None => (self.stand_in.view, None),
};
Self::write_image_descriptor(
device,
@@ -1903,26 +1507,15 @@ impl RtStage {
view,
self.image_sampler,
);
- let params = RtParams {
- inv_mvp: camera.inv_mvp,
- width: self.output_size.0,
- height: self.output_size.1,
- sample_index: self.sample_index,
- max_bounces: MAX_BOUNCES,
- spp: self.spp,
- _pad: [0; 3],
- img_origin,
- img_u,
- img_v,
- background: match self.background {
- Some([r, g, b]) => [r, g, b, 1.0],
- None => [0.0; 4],
- },
- sun_dir: v4(self.environment.sun_unit()),
- sun_color: v4(self.environment.sun_color),
- sky_zenith: v4(self.environment.sky_zenith),
- sky_nadir: v4(self.environment.sky_nadir),
- };
+ let params = rt_params(
+ camera,
+ self.output_size,
+ self.sample_index,
+ self.spp,
+ param_image,
+ self.background,
+ &self.environment,
+ );
let frame = &mut self.frames[frame_index];
frame.uniforms.allocation.as_mut().unwrap().mapped_slice_mut().unwrap()
[..std::mem::size_of::<RtParams>()]
@@ -2512,229 +2105,14 @@ impl Drop for RtOffscreen {
mod tests {
use super::*;
- // CPU mirror of the shader's traversal, for parity testing.
- fn intersect_tri_cpu(ro: [f32; 3], rd: [f32; 3], t: &RtTriangle, t_limit: f32) -> f32 {
- let sub = |a: [f32; 3], b: [f32; 3]| [a[0] - b[0], a[1] - b[1], a[2] - b[2]];
- let cross = |a: [f32; 3], b: [f32; 3]| {
- [
- a[1] * b[2] - a[2] * b[1],
- a[2] * b[0] - a[0] * b[2],
- a[0] * b[1] - a[1] * b[0],
- ]
- };
- let dot = |a: [f32; 3], b: [f32; 3]| a[0] * b[0] + a[1] * b[1] + a[2] * b[2];
- let e1 = sub(t.p1, t.p0);
- let e2 = sub(t.p2, t.p0);
- let h = cross(rd, e2);
- let a = dot(e1, h);
- if a.abs() < 1e-8 {
- return 1e30;
- }
- let f = 1.0 / a;
- let s = sub(ro, t.p0);
- let u = f * dot(s, h);
- if !(0.0..=1.0).contains(&u) {
- return 1e30;
- }
- let q = cross(s, e1);
- let v = f * dot(rd, q);
- if v < 0.0 || u + v > 1.0 {
- return 1e30;
- }
- let tt = f * dot(e2, q);
- if tt > 1e-4 && tt < t_limit {
- return tt;
- }
- 1e30
- }
-
- fn traverse_bvh_cpu(
- nodes: &[GpuBvhNode],
- tris: &[RtTriangle],
- ro: [f32; 3],
- rd: [f32; 3],
- ) -> (f32, Option<usize>) {
- if nodes.is_empty() {
- return (1e30, None);
- }
- let inv = [1.0 / rd[0], 1.0 / rd[1], 1.0 / rd[2]];
- let hit_aabb = |min: [f32; 3], max: [f32; 3], t_limit: f32| -> bool {
- let mut tn = f32::NEG_INFINITY;
- let mut tf = f32::INFINITY;
- for a in 0..3 {
- let t1 = (min[a] - ro[a]) * inv[a];
- let t2 = (max[a] - ro[a]) * inv[a];
- tn = tn.max(t1.min(t2));
- tf = tf.min(t1.max(t2));
- }
- tf >= tn.max(0.0) && tn < t_limit
- };
- let mut best = 1e30f32;
- let mut best_tri = None;
- let mut stack = vec![0u32];
- while let Some(idx) = stack.pop() {
- let node = &nodes[idx as usize];
- if !hit_aabb(node.min, node.max, best) {
- continue;
- }
- if node.count > 0 {
- for i in node.left_first..node.left_first + node.count {
- let t = intersect_tri_cpu(ro, rd, &tris[i as usize], best);
- if t < best {
- best = t;
- best_tri = Some(i as usize);
- }
- }
- } else {
- stack.push(node.left_first);
- stack.push(node.left_first + 1);
- }
- }
- (best, best_tri)
- }
-
- fn brute_force(tris: &[RtTriangle], ro: [f32; 3], rd: [f32; 3]) -> (f32, Option<usize>) {
- let mut best = 1e30f32;
- let mut best_tri = None;
- for (i, t) in tris.iter().enumerate() {
- let tt = intersect_tri_cpu(ro, rd, t, best);
- if tt < best {
- best = tt;
- best_tri = Some(i);
- }
- }
- (best, best_tri)
- }
-
- // Deterministic LCG so the test needs no rand dependency.
- struct Lcg(u64);
- impl Lcg {
- fn next_f32(&mut self) -> f32 {
- self.0 = self.0.wrapping_mul(6364136223846793005).wrapping_add(1442695040888963407);
- ((self.0 >> 33) as f32) / (u32::MAX >> 1) as f32
- }
- fn point(&mut self, scale: f32) -> [f32; 3] {
- [
- (self.next_f32() - 0.5) * scale,
- (self.next_f32() - 0.5) * scale,
- (self.next_f32() - 0.5) * scale,
- ]
- }
- }
-
- fn random_scene(n: usize, seed: u64) -> Vec<RtTriangle> {
- let mut rng = Lcg(seed);
- (0..n)
- .map(|i| {
- let c = rng.point(20.0);
- let jitter = |rng: &mut Lcg, c: [f32; 3]| {
- let d = rng.point(2.0);
- [c[0] + d[0], c[1] + d[1], c[2] + d[2]]
- };
- RtTriangle {
- p0: jitter(&mut rng, c),
- p1: jitter(&mut rng, c),
- p2: jitter(&mut rng, c),
- material: (i % 5) as u32,
- }
- })
- .collect()
- }
-
- #[test]
- fn test_bvh_matches_brute_force() {
- let mut tris = random_scene(500, 42);
- let nodes = build_bvh(&mut tris);
- assert!(!nodes.is_empty());
- let mut rng = Lcg(7);
- let mut hits = 0;
- for _ in 0..200 {
- let ro = rng.point(40.0);
- let target = rng.point(10.0);
- let d = [target[0] - ro[0], target[1] - ro[1], target[2] - ro[2]];
- let len = (d[0] * d[0] + d[1] * d[1] + d[2] * d[2]).sqrt().max(1e-6);
- let rd = [d[0] / len, d[1] / len, d[2] / len];
- let (t_bvh, tri_bvh) = traverse_bvh_cpu(&nodes, &tris, ro, rd);
- let (t_ref, tri_ref) = brute_force(&tris, ro, rd);
- assert_eq!(tri_bvh, tri_ref, "different triangle hit");
- assert!((t_bvh - t_ref).abs() < 1e-4, "t mismatch: {t_bvh} vs {t_ref}");
- if tri_bvh.is_some() {
- hits += 1;
- }
- }
- assert!(hits > 20, "test rays barely hit the scene ({hits}/200)");
- }
-
- #[test]
- fn test_bvh_leaf_ranges_cover_all_triangles() {
- let mut tris = random_scene(300, 9);
- let nodes = build_bvh(&mut tris);
- let mut seen = vec![false; tris.len()];
- for node in &nodes {
- if node.count > 0 {
- for i in node.left_first..node.left_first + node.count {
- assert!(!seen[i as usize], "triangle {i} in two leaves");
- seen[i as usize] = true;
- }
- }
- }
- assert!(seen.iter().all(|&s| s), "not every triangle is in a leaf");
- }
-
- #[test]
- fn test_bvh_degenerate_identical_centroids() {
- // All triangles share one centroid: SAH can't split, the median
- // fallback must still terminate and cover everything.
- let tri = RtTriangle {
- p0: [0.0, 0.0, 0.0],
- p1: [1.0, 0.0, 0.0],
- p2: [0.0, 1.0, 0.0],
- material: 0,
- };
- let mut tris = vec![tri; 100];
- let nodes = build_bvh(&mut tris);
- let covered: u32 = nodes.iter().filter(|n| n.count > 0).map(|n| n.count).sum();
- assert_eq!(covered, 100);
- let (t, hit) = traverse_bvh_cpu(&nodes, &tris, [0.2, 0.2, -5.0], [0.0, 0.0, 1.0]);
- assert!(hit.is_some());
- assert!((t - 5.0).abs() < 1e-3);
- }
-
- #[test]
- fn test_bvh_empty_and_single() {
- let mut empty: Vec<RtTriangle> = Vec::new();
- assert!(build_bvh(&mut empty).is_empty());
-
- let mut single = vec![RtTriangle {
- p0: [-1.0, -1.0, 0.0],
- p1: [1.0, -1.0, 0.0],
- p2: [0.0, 1.0, 0.0],
- material: 3,
- }];
- let nodes = build_bvh(&mut single);
- assert_eq!(nodes.len(), 1);
- assert_eq!(nodes[0].count, 1);
- let (t, hit) = traverse_bvh_cpu(&nodes, &single, [0.0, 0.0, -3.0], [0.0, 0.0, 1.0]);
- assert_eq!(hit, Some(0));
- assert!((t - 3.0).abs() < 1e-4);
- }
-
#[test]
fn test_rt_shaders_compile() {
// naga parse + validate + SPIR-V write for both tiers; panics on failure.
- let tier1 = compile_wgsl(&format!(
- "{}\n{}",
- include_str!("rt_common.wgsl"),
- include_str!("rt_bvh.wgsl")
- ));
+ let tier1 = compile_wgsl(&crate::draw::shaders::rt_bvh_source());
assert!(!tier1.is_empty());
- let tier2 = compile_wgsl_ray_query(&format!(
- "{}\n{}",
- include_str!("rt_common.wgsl"),
- include_str!("rt_query.wgsl")
- ));
+ let tier2 = compile_wgsl_ray_query(&format!("{}\n{}", crate::draw::shaders::RT_COMMON, crate::draw::shaders::RT_QUERY));
assert!(!tier2.is_empty());
- let denoise = compile_wgsl(include_str!("rt_denoise.wgsl"));
+ let denoise = compile_wgsl(crate::draw::shaders::RT_DENOISE);
assert!(!denoise.is_empty());
}
diff --git a/src/web/mod.rs b/src/web/mod.rs
index 97041ac..d6871da 100644
--- a/src/web/mod.rs
+++ b/src/web/mod.rs
@@ -11,6 +11,7 @@
mod compute;
mod renderer;
+mod rt;
mod scene;
mod shell;
diff --git a/src/web/renderer.rs b/src/web/renderer.rs
index c4df5c6..7ebae23 100644
--- a/src/web/renderer.rs
+++ b/src/web/renderer.rs
@@ -26,7 +26,9 @@
use std::collections::HashMap;
+use super::rt::WebRt;
use super::scene::WebScene;
+use crate::draw::rt::{RtCamera, RtEnvironment, RtImage, RtMaterial, RtTriangle};
use crate::draw::scene::{MeshId, SceneDraw, SceneImage, Stage3D, Vertex3D};
use wasm_bindgen::{JsCast, JsValue};
@@ -171,6 +173,11 @@ pub struct WebRenderer {
/// the backdrop) — what the 2D pass binds while a scene is shown.
scene: WebScene,
scene_group_0: Option<GpuBindGroup>,
+ /// The path tracer, made by the first `set_rt_scene`; the background and
+ /// environment it is handed at each `stage_rt`, as on Vulkan.
+ rt: Option<WebRt>,
+ rt_background: Option<[f32; 3]>,
+ rt_environment: RtEnvironment,
blocks: Growable,
group_1: GpuBindGroup,
vertices: Growable,
@@ -444,6 +451,9 @@ impl WebRenderer {
snapshot: None,
scene,
scene_group_0: None,
+ rt: None,
+ rt_background: None,
+ rt_environment: RtEnvironment::default(),
blocks,
group_1,
vertices,
@@ -610,7 +620,8 @@ impl WebRenderer {
let (w, h) = (target.width(), target.height());
// The scene's backdrop follows the canvas; a resized one holds no
// scene until the next is drawn into it, as on Vulkan.
- if self.scene.has_staged() || self.scene.target.is_some() {
+ let rt_staged = self.rt.as_ref().is_some_and(|rt| rt.staged());
+ if self.scene.has_staged() || rt_staged || self.scene.target.is_some() {
let had = self.scene.target.as_ref().map(|t| (t.width, t.height));
self.scene.fit(&self.device, w, h)?;
if had != Some((w, h)) {
@@ -689,6 +700,16 @@ impl WebRenderer {
let encoder = self.device.create_command_encoder();
let images_for_scene = &self.images;
self.scene.record(&self.device, &self.queue, &encoder, &|id| images_for_scene.get(&id).map(|i| i.group.clone()))?;
+ // The traced pane, into the same backdrop after the raster scene.
+ if let (Some(rt), Some(t)) = (self.rt.as_mut(), self.scene.target.as_ref()) {
+ let image_view = |id: u32| {
+ let img = images_for_scene.get(&id)?;
+ Some((img.texture.create_view().ok()?, img.width, img.height))
+ };
+ if rt.record(&self.device, &self.queue, &encoder, &t.view, (t.width, t.height), &image_view)? {
+ self.scene.backdrop_valid = true;
+ }
+ }
let base_group_0 = match (&self.scene_group_0, &self.scene.target) {
(Some(group), Some(t)) if self.scene.backdrop_valid => {
encoder.copy_texture_to_texture_with_gpu_extent_3d_dict(
@@ -861,4 +882,37 @@ impl Stage3D for WebRenderer {
self.scene.light = v.normalize().to_array();
}
}
+ fn set_rt_scene_with_image(&mut self, triangles: &[RtTriangle], materials: &[RtMaterial], image: Option<RtImage>) {
+ if self.rt.is_none() {
+ match WebRt::new(&self.device, self.view_format) {
+ Ok(rt) => self.rt = Some(rt),
+ Err(e) => {
+ web_sys::console::error_2(&"cce-ui: no path tracer:".into(), &e);
+ return;
+ }
+ }
+ }
+ let rt = self.rt.as_mut().unwrap();
+ if let Err(e) = rt.set_scene(&self.device, &self.queue, triangles, materials, image) {
+ web_sys::console::error_2(&"cce-ui: the traced scene was not uploaded:".into(), &e);
+ }
+ }
+ fn set_rt_environment(&mut self, environment: RtEnvironment) {
+ self.rt_environment = environment;
+ }
+ fn set_rt_background(&mut self, color: Option<[f32; 3]>) {
+ self.rt_background = color;
+ }
+ fn stage_rt(&mut self, pane: (u32, u32, u32, u32), camera: RtCamera) {
+ if let Some(rt) = self.rt.as_mut() {
+ rt.set_background(self.rt_background);
+ rt.set_environment(self.rt_environment);
+ if let Err(e) = rt.stage(&self.device, pane, camera) {
+ web_sys::console::error_2(&"cce-ui: the traced pane was not staged:".into(), &e);
+ }
+ }
+ }
+ fn rt_accumulating(&self) -> bool {
+ self.rt.as_ref().is_some_and(|rt| rt.accumulating())
+ }
}
diff --git a/src/web/rt.rs b/src/web/rt.rs
new file mode 100644
index 0000000..e0bf597
--- /dev/null
+++ b/src/web/rt.rs
@@ -0,0 +1,469 @@
+//! The path tracer on WebGPU: the Vulkan `RtStage`'s compute tier, ported.
+//! The same shaders (`draw::shaders::rt_bvh_source`, `RT_DENOISE`), the
+//! same scene packing, BVH and parameter blocks (`draw::rt`), the same
+//! rules for when the accumulation restarts — a camera, pane size, scene,
+//! background or environment change — and the same frame: one sample a
+//! frame added into the running mean, three à-trous denoise iterations over
+//! it, the result put into the backdrop's pane region in place of the
+//! raster scene, a pane that moved clearing the backdrop once first.
+//!
+//! What WebGPU makes different:
+//!
+//! - **The compute tier only.** The hardware ray-query tier needs
+//! `VK_KHR_ray_query`; WebGPU has no ray tracing. A BVH traversed in
+//! compute is the tier every Vulkan device without RT cores runs too.
+//! - **The blit is a draw.** Vulkan blits the tracer's `rgba8unorm` image
+//! into the sRGB backdrop, converting as it copies; WebGPU copies only
+//! between formats that differ in sRGB-ness at most, so a small render
+//! pass loads each texel and writes it through the backdrop's sRGB view —
+//! the same conversion (the unorm value taken as linear and encoded).
+//! - **No frames in flight to juggle.** One parameter buffer and one bind
+//! group per dispatch, written before the frame's single submission.
+
+use wasm_bindgen::JsValue;
+use web_sys::{
+ gpu_buffer_usage as buffer_usage, gpu_shader_stage as shader_stage, gpu_texture_usage as texture_usage, GpuBindGroup,
+ GpuBindGroupDescriptor, GpuBindGroupEntry, GpuBindGroupLayout, GpuBindGroupLayoutDescriptor, GpuBindGroupLayoutEntry,
+ GpuBuffer, GpuBufferBinding, GpuBufferBindingLayout, GpuBufferBindingType, GpuColorTargetState, GpuCommandEncoder,
+ GpuComputePassDescriptor, GpuComputePipeline, GpuComputePipelineDescriptor, GpuDevice, GpuFragmentState, GpuLoadOp,
+ GpuPipelineLayoutDescriptor, GpuProgrammableStage, GpuQueue, GpuRenderPassColorAttachment, GpuRenderPassDescriptor,
+ GpuRenderPipeline, GpuRenderPipelineDescriptor, GpuSampler, GpuSamplerBindingLayout, GpuSamplerBindingType,
+ GpuStorageTextureAccess, GpuStorageTextureBindingLayout, GpuStoreOp, GpuTexture, GpuTextureBindingLayout,
+ GpuTextureFormat, GpuTextureSampleType, GpuTextureView, GpuTextureViewDimension, GpuVertexState,
+};
+
+use super::renderer::{shader_module, texture, whole_view};
+use crate::draw::rt::{
+ denoise_params, pack_scene, rt_params, DenoiseParams, ParamImage, RtCamera, RtEnvironment, RtImage, RtMaterial,
+ RtParams, RtTriangle, DENOISE_ITERATIONS, MAX_SAMPLES, WORKGROUP,
+};
+use crate::draw::shaders::{rt_bvh_source, RT_DENOISE};
+
+/// Each denoise iteration's block sits at a multiple of this.
+const STRIDE: u32 = 256;
+
+/// Loads the tracer's image at the fragment's place in the pane and writes
+/// it through the backdrop's sRGB view: Vulkan's UNORM-to-sRGB blit.
+const BLIT: &str = "
+struct Blit { origin: vec2<i32>, _pad: vec2<i32> }
+@group(0) @binding(0) var src: texture_2d<f32>;
+@group(0) @binding(1) var<uniform> blit: Blit;
+@vertex fn vs_main(@builtin(vertex_index) i: u32) -> @builtin(position) vec4<f32> {
+ let p = vec2<f32>(f32((i << 1u) & 2u), f32(i & 2u));
+ return vec4<f32>(p * 2.0 - 1.0, 0.0, 1.0);
+}
+@fragment fn fs_main(@builtin(position) pos: vec4<f32>) -> @location(0) vec4<f32> {
+ return textureLoad(src, vec2<i32>(pos.xy) - blit.origin, 0);
+}";
+
+/// The pane-sized targets: the running sum, the denoiser's features and
+/// ping-pong, and the image the result is written to.
+struct Targets {
+ accum: GpuBuffer,
+ features: GpuBuffer,
+ ping: GpuBuffer,
+ pong: GpuBuffer,
+ output: GpuTexture,
+ output_view: GpuTextureView,
+ /// Denoise iterations 0 and 2 (src pong, dst ping), and 1 (src ping, dst pong).
+ denoise_b: GpuBindGroup,
+ denoise_a: GpuBindGroup,
+ blit: GpuBindGroup,
+ size: (u32, u32),
+}
+
+/// The scene the buffers hold, and the image standing in it.
+struct Scene {
+ nodes: GpuBuffer,
+ tris: GpuBuffer,
+ materials: GpuBuffer,
+ tri_count: u32,
+ image: Option<RtImage>,
+}
+
+pub(crate) struct WebRt {
+ trace: GpuComputePipeline,
+ trace_layout: GpuBindGroupLayout,
+ denoise: GpuComputePipeline,
+ denoise_layout: GpuBindGroupLayout,
+ blit: GpuRenderPipeline,
+ blit_layout: GpuBindGroupLayout,
+ params: GpuBuffer,
+ denoise_params: GpuBuffer,
+ blit_params: GpuBuffer,
+ stand_in: GpuTextureView,
+ sampler: GpuSampler,
+ scene: Option<Scene>,
+ targets: Option<Targets>,
+
+ pane: (u32, u32, u32, u32),
+ pane_moved: bool,
+ camera: Option<RtCamera>,
+ sample_index: u32,
+ spp: u32,
+ staged: bool,
+ background: Option<[f32; 3]>,
+ environment: RtEnvironment,
+}
+
+fn buffer_entry(binding: u32, ty: GpuBufferBindingType, dynamic: bool) -> GpuBindGroupLayoutEntry {
+ let entry = GpuBindGroupLayoutEntry::new(binding, shader_stage::COMPUTE);
+ let layout = GpuBufferBindingLayout::new();
+ layout.set_type(ty);
+ layout.set_has_dynamic_offset(dynamic);
+ entry.set_buffer(&layout);
+ entry
+}
+
+fn storage_texture_entry(binding: u32) -> GpuBindGroupLayoutEntry {
+ let entry = GpuBindGroupLayoutEntry::new(binding, shader_stage::COMPUTE);
+ let layout = GpuStorageTextureBindingLayout::new(GpuTextureFormat::Rgba8unorm);
+ layout.set_access(GpuStorageTextureAccess::WriteOnly);
+ layout.set_view_dimension(GpuTextureViewDimension::N2d);
+ entry.set_storage_texture(&layout);
+ entry
+}
+
+fn texture_entry(binding: u32, visibility: u32, sample: GpuTextureSampleType) -> GpuBindGroupLayoutEntry {
+ let entry = GpuBindGroupLayoutEntry::new(binding, visibility);
+ let layout = GpuTextureBindingLayout::new();
+ layout.set_sample_type(sample);
+ layout.set_view_dimension(GpuTextureViewDimension::N2d);
+ entry.set_texture(&layout);
+ entry
+}
+
+fn whole(binding: u32, buffer: &GpuBuffer) -> GpuBindGroupEntry {
+ GpuBindGroupEntry::new_with_gpu_buffer_binding(binding, &GpuBufferBinding::new(buffer))
+}
+
+fn sized(binding: u32, buffer: &GpuBuffer, size: u32) -> GpuBindGroupEntry {
+ let b = GpuBufferBinding::new(buffer);
+ b.set_size(size);
+ GpuBindGroupEntry::new_with_gpu_buffer_binding(binding, &b)
+}
+
+fn buffer(device: &GpuDevice, size: u32, usage: u32, label: &str) -> Result<GpuBuffer, JsValue> {
+ let desc = web_sys::GpuBufferDescriptor::new(size.max(16).next_multiple_of(4), usage);
+ desc.set_label(label);
+ device.create_buffer(&desc)
+}
+
+fn compute_pipeline(device: &GpuDevice, layout: &GpuBindGroupLayout, code: &str, entry: &str, label: &str) -> GpuComputePipeline {
+ let pipeline_layout = device.create_pipeline_layout(&GpuPipelineLayoutDescriptor::new(&[js_sys::JsOption::wrap(layout.clone())]));
+ let stage = GpuProgrammableStage::new(&shader_module(device, code, label));
+ stage.set_entry_point(entry);
+ let desc = GpuComputePipelineDescriptor::new(&pipeline_layout, &stage);
+ desc.set_label(label);
+ device.create_compute_pipeline(&desc)
+}
+
+impl WebRt {
+ /// The tracer's pipelines; `backdrop_format` is what the blit writes.
+ pub(crate) fn new(device: &GpuDevice, backdrop_format: GpuTextureFormat) -> Result<Self, JsValue> {
+ use GpuBufferBindingType::{ReadOnlyStorage, Storage, Uniform};
+ let sampler_entry = {
+ let entry = GpuBindGroupLayoutEntry::new(8, shader_stage::COMPUTE);
+ let layout = GpuSamplerBindingLayout::new();
+ layout.set_type(GpuSamplerBindingType::Filtering);
+ entry.set_sampler(&layout);
+ entry
+ };
+ let trace_layout = device.create_bind_group_layout(&GpuBindGroupLayoutDescriptor::new(&[
+ buffer_entry(0, Uniform, false),
+ buffer_entry(1, ReadOnlyStorage, false),
+ buffer_entry(2, ReadOnlyStorage, false),
+ buffer_entry(3, ReadOnlyStorage, false),
+ buffer_entry(4, Storage, false),
+ storage_texture_entry(5),
+ buffer_entry(6, Storage, false),
+ texture_entry(7, shader_stage::COMPUTE, GpuTextureSampleType::Float),
+ sampler_entry,
+ ]))?;
+ let trace = compute_pipeline(device, &trace_layout, &rt_bvh_source(), "cs_main", "rt-trace");
+ let denoise_layout = device.create_bind_group_layout(&GpuBindGroupLayoutDescriptor::new(&[
+ buffer_entry(0, Uniform, true),
+ buffer_entry(1, ReadOnlyStorage, false),
+ buffer_entry(2, ReadOnlyStorage, false),
+ buffer_entry(3, ReadOnlyStorage, false),
+ buffer_entry(4, Storage, false),
+ storage_texture_entry(5),
+ ]))?;
+ let denoise = compute_pipeline(device, &denoise_layout, RT_DENOISE, "cs_denoise", "rt-denoise");
+
+ let blit_layout = device.create_bind_group_layout(&GpuBindGroupLayoutDescriptor::new(&[
+ texture_entry(0, shader_stage::FRAGMENT, GpuTextureSampleType::UnfilterableFloat),
+ {
+ let entry = GpuBindGroupLayoutEntry::new(1, shader_stage::FRAGMENT);
+ let layout = GpuBufferBindingLayout::new();
+ layout.set_type(GpuBufferBindingType::Uniform);
+ entry.set_buffer(&layout);
+ entry
+ },
+ ]))?;
+ let blit_module = shader_module(device, BLIT, "rt-blit");
+ let vertex = GpuVertexState::new(&blit_module);
+ vertex.set_entry_point("vs_main");
+ let targets = [js_sys::JsOption::wrap(GpuColorTargetState::new(backdrop_format))];
+ let fragment = GpuFragmentState::new(&blit_module, &targets);
+ fragment.set_entry_point("fs_main");
+ let desc = GpuRenderPipelineDescriptor::new(
+ &device.create_pipeline_layout(&GpuPipelineLayoutDescriptor::new(&[js_sys::JsOption::wrap(blit_layout.clone())])),
+ &vertex,
+ );
+ desc.set_fragment(&fragment);
+ desc.set_label("rt-blit");
+ let blit = device.create_render_pipeline(&desc)?;
+
+ let params = buffer(device, std::mem::size_of::<RtParams>() as u32, buffer_usage::UNIFORM | buffer_usage::COPY_DST, "rt-params")?;
+ let denoise_params =
+ buffer(device, STRIDE * DENOISE_ITERATIONS as u32, buffer_usage::UNIFORM | buffer_usage::COPY_DST, "rt-denoise-params")?;
+ let blit_params = buffer(device, 16, buffer_usage::UNIFORM | buffer_usage::COPY_DST, "rt-blit-params")?;
+ // The image binding's stand-in while the scene has none: one clear texel.
+ let stand_in = texture(device, GpuTextureFormat::Rgba8unormSrgb, 1, 1, texture_usage::TEXTURE_BINDING | texture_usage::COPY_DST, "rt-stand-in")?;
+ let sampler = {
+ let desc = web_sys::GpuSamplerDescriptor::new();
+ desc.set_mag_filter(web_sys::GpuFilterMode::Linear);
+ desc.set_min_filter(web_sys::GpuFilterMode::Linear);
+ desc.set_mipmap_filter(web_sys::GpuMipmapFilterMode::Linear);
+ device.create_sampler_with_descriptor(&desc)
+ };
+ Ok(Self {
+ trace,
+ trace_layout,
+ denoise,
+ denoise_layout,
+ blit,
+ blit_layout,
+ params,
+ denoise_params,
+ blit_params,
+ stand_in: whole_view(&stand_in)?,
+ sampler,
+ scene: None,
+ targets: None,
+ pane: (0, 0, 0, 0),
+ pane_moved: false,
+ camera: None,
+ sample_index: 0,
+ spp: 1,
+ staged: false,
+ background: None,
+ environment: RtEnvironment::default(),
+ })
+ }
+
+ /// Replace the scene: packed and its BVH built on the CPU, uploaded.
+ pub(crate) fn set_scene(
+ &mut self,
+ device: &GpuDevice,
+ queue: &GpuQueue,
+ triangles: &[RtTriangle],
+ materials: &[RtMaterial],
+ image: Option<RtImage>,
+ ) -> Result<(), JsValue> {
+ let packed = pack_scene(triangles, materials, image.map(|i| i.corners), true);
+ let upload = |bytes: &[u8], label: &str| -> Result<GpuBuffer, JsValue> {
+ let b = buffer(device, bytes.len() as u32, buffer_usage::STORAGE | buffer_usage::COPY_DST, label)?;
+ if !bytes.is_empty() {
+ queue.write_buffer_with_u32_and_u8_slice(&b, 0, bytes)?;
+ }
+ Ok(b)
+ };
+ if let Some(old) = self.scene.take() {
+ for b in [old.nodes, old.tris, old.materials] {
+ b.destroy();
+ }
+ }
+ self.scene = Some(Scene {
+ nodes: upload(bytemuck::cast_slice(&packed.nodes), "rt-nodes")?,
+ tris: upload(bytemuck::cast_slice(&packed.tris), "rt-tris")?,
+ materials: upload(bytemuck::cast_slice(&packed.materials), "rt-materials")?,
+ tri_count: packed.tris.len() as u32,
+ image,
+ });
+ self.sample_index = 0;
+ Ok(())
+ }
+
+ pub(crate) fn set_background(&mut self, background: Option<[f32; 3]>) {
+ if self.background != background {
+ self.background = background;
+ self.sample_index = 0;
+ }
+ }
+
+ pub(crate) fn set_environment(&mut self, environment: RtEnvironment) {
+ if self.environment != environment {
+ self.environment = environment;
+ self.sample_index = 0;
+ }
+ }
+
+ /// Stage a frame for `pane` (physical px): the targets follow its size,
+ /// and a camera or size change restarts the accumulation; a pane that
+ /// moved clears the backdrop once before the next blit.
+ pub(crate) fn stage(&mut self, device: &GpuDevice, pane: (u32, u32, u32, u32), camera: RtCamera) -> Result<(), JsValue> {
+ let (_, _, w, h) = pane;
+ if w == 0 || h == 0 {
+ self.staged = false;
+ return Ok(());
+ }
+ if self.targets.as_ref().map(|t| t.size) != Some((w, h)) {
+ self.recreate_targets(device, w, h)?;
+ self.sample_index = 0;
+ }
+ if pane != self.pane && self.pane != (0, 0, 0, 0) {
+ self.pane_moved = true;
+ }
+ if self.camera != Some(camera) {
+ self.camera = Some(camera);
+ self.sample_index = 0;
+ }
+ self.pane = pane;
+ self.staged = true;
+ Ok(())
+ }
+
+ pub(crate) fn staged(&self) -> bool {
+ self.staged
+ }
+
+ /// True while another dispatch would still refine the image.
+ pub(crate) fn accumulating(&self) -> bool {
+ self.scene.as_ref().is_some_and(|s| s.tri_count > 0) && self.sample_index < MAX_SAMPLES
+ }
+
+ fn recreate_targets(&mut self, device: &GpuDevice, w: u32, h: u32) -> Result<(), JsValue> {
+ if let Some(old) = self.targets.take() {
+ for b in [old.accum, old.features, old.ping, old.pong] {
+ b.destroy();
+ }
+ old.output.destroy();
+ }
+ let px = w * h;
+ let storage = buffer_usage::STORAGE;
+ let accum = buffer(device, px * 16, storage, "rt-accum")?;
+ let features = buffer(device, px * 32, storage, "rt-features")?;
+ let ping = buffer(device, px * 16, storage, "rt-denoise-ping")?;
+ let pong = buffer(device, px * 16, storage, "rt-denoise-pong")?;
+ let output = texture(
+ device,
+ GpuTextureFormat::Rgba8unorm,
+ w,
+ h,
+ texture_usage::STORAGE_BINDING | texture_usage::TEXTURE_BINDING,
+ "rt-output",
+ )?;
+ let output_view = whole_view(&output)?;
+ let denoise_group = |src: &GpuBuffer, dst: &GpuBuffer| {
+ let entries = [
+ sized(0, &self.denoise_params, std::mem::size_of::<DenoiseParams>() as u32),
+ whole(1, &accum),
+ whole(2, &features),
+ whole(3, src),
+ whole(4, dst),
+ GpuBindGroupEntry::new_with_gpu_texture_view(5, &output_view),
+ ];
+ device.create_bind_group(&GpuBindGroupDescriptor::new(&entries, &self.denoise_layout))
+ };
+ let denoise_b = denoise_group(&pong, &ping);
+ let denoise_a = denoise_group(&ping, &pong);
+ let blit = device.create_bind_group(&GpuBindGroupDescriptor::new(
+ &[GpuBindGroupEntry::new_with_gpu_texture_view(0, &output_view), whole(1, &self.blit_params)],
+ &self.blit_layout,
+ ));
+ self.targets = Some(Targets { accum, features, ping, pong, output, output_view, denoise_b, denoise_a, blit, size: (w, h) });
+ Ok(())
+ }
+
+ /// Record one accumulation dispatch, the denoise and the blit into
+ /// `backdrop` (the scene's, `backdrop_size` px). False when there is
+ /// nothing to do (not staged, an empty scene, converged): the backdrop
+ /// then keeps what it has. `image_view` looks the scene's image up in
+ /// the 2D pass's table: its view and size, if it has landed.
+ pub(crate) fn record(
+ &mut self,
+ device: &GpuDevice,
+ queue: &GpuQueue,
+ encoder: &GpuCommandEncoder,
+ backdrop: &GpuTextureView,
+ backdrop_size: (u32, u32),
+ image_view: &dyn Fn(u32) -> Option<(GpuTextureView, u32, u32)>,
+ ) -> Result<bool, JsValue> {
+ let ready = self.staged && self.scene.as_ref().is_some_and(|s| s.tri_count > 0) && self.targets.is_some();
+ self.staged = false;
+ if !ready || self.sample_index >= MAX_SAMPLES {
+ return Ok(false);
+ }
+ let (Some(scene), Some(t), Some(camera)) = (&self.scene, &self.targets, self.camera) else { return Ok(false) };
+ let (w, h) = t.size;
+
+ // The parameter blocks: the scene's image as it is NOW (its upload
+ // may land after the scene was set), or none, which lets rays through.
+ let bound = scene.image.and_then(|i| image_view(i.image).map(|(view, iw, ih)| (view, i, iw, ih)));
+ let (view, param_image) = match bound {
+ Some((view, i, iw, ih)) => (view, Some(ParamImage { width: iw, height: ih, corners: i.corners, opacity: i.opacity })),
+ None => (self.stand_in.clone(), None),
+ };
+ let params = rt_params(camera, (w, h), self.sample_index, self.spp, param_image, self.background, &self.environment);
+ queue.write_buffer_with_u32_and_u8_slice(&self.params, 0, bytemuck::bytes_of(¶ms))?;
+ let mut blocks = vec![0u8; STRIDE as usize * DENOISE_ITERATIONS];
+ for (i, p) in denoise_params(w, h, self.sample_index + self.spp).iter().enumerate() {
+ let at = i * STRIDE as usize;
+ blocks[at..at + std::mem::size_of::<DenoiseParams>()].copy_from_slice(bytemuck::bytes_of(p));
+ }
+ queue.write_buffer_with_u32_and_u8_slice(&self.denoise_params, 0, &blocks)?;
+ let (px, py, _, _) = self.pane;
+ let dst_x = px.min(backdrop_size.0);
+ let dst_y = py.min(backdrop_size.1);
+ let (bw, bh) = (w.min(backdrop_size.0 - dst_x), h.min(backdrop_size.1 - dst_y));
+ queue.write_buffer_with_u32_and_u8_slice(&self.blit_params, 0, bytemuck::cast_slice(&[dst_x as i32, dst_y as i32, 0, 0]))?;
+
+ let trace_group = device.create_bind_group(&GpuBindGroupDescriptor::new(
+ &[
+ whole(0, &self.params),
+ whole(1, &scene.nodes),
+ whole(2, &scene.tris),
+ whole(3, &scene.materials),
+ whole(4, &t.accum),
+ GpuBindGroupEntry::new_with_gpu_texture_view(5, &t.output_view),
+ whole(6, &t.features),
+ GpuBindGroupEntry::new_with_gpu_texture_view(7, &view),
+ GpuBindGroupEntry::new(8, &self.sampler),
+ ],
+ &self.trace_layout,
+ ));
+ let groups = (w.div_ceil(WORKGROUP), h.div_ceil(WORKGROUP));
+ let pass = encoder.begin_compute_pass_with_descriptor(&GpuComputePassDescriptor::new());
+ pass.set_pipeline(&self.trace);
+ pass.set_bind_group(0, Some(&trace_group));
+ pass.dispatch_workgroups_with_workgroup_count_y(groups.0, groups.1);
+ // À-trous iterations: 0 and 2 write ping, 1 writes pong; the last
+ // rewrites the output image.
+ pass.set_pipeline(&self.denoise);
+ for i in 0..DENOISE_ITERATIONS {
+ let group = if i % 2 == 0 { &t.denoise_b } else { &t.denoise_a };
+ pass.set_bind_group_with_u32_slice_and_u32_and_dynamic_offsets_data_length(0, Some(group), &[i as u32 * STRIDE], 0, 1)?;
+ pass.dispatch_workgroups_with_workgroup_count_y(groups.0, groups.1);
+ }
+ pass.end();
+
+ // Into the backdrop's pane region (clearing the whole backdrop first
+ // when the pane moved: stale pixels sit outside the new region).
+ let load = if std::mem::take(&mut self.pane_moved) { GpuLoadOp::Clear } else { GpuLoadOp::Load };
+ let color = GpuRenderPassColorAttachment::new_with_gpu_texture_view(load, GpuStoreOp::Store, backdrop);
+ color.set_clear_value(&[0.0, 0.0, 0.0, 0.0].map(js_sys::Number::from));
+ let pass = encoder.begin_render_pass(&GpuRenderPassDescriptor::new(&[js_sys::JsOption::wrap(color)]))?;
+ if bw > 0 && bh > 0 {
+ pass.set_viewport(dst_x as f32, dst_y as f32, bw as f32, bh as f32, 0.0, 1.0);
+ pass.set_scissor_rect(dst_x, dst_y, bw, bh);
+ pass.set_pipeline(&self.blit);
+ pass.set_bind_group(0, Some(&t.blit));
+ pass.draw(3);
+ }
+ pass.end();
+ self.sample_index += self.spp;
+ Ok(true)
+ }
+}