Repository navigation
Add small compute examples illustrating new WGSL primitives for AI #350
Description
Activity
CC @dneto0
For dp4a:
- They're good for ML image processing: segmentation, object identification. But it's involved to put an ML model in the samples.
- Maybe we can do something really simple, like sobel filter on a grayscale image.
Reacted by Kai Ninomiyadp4a accelerates matrix multiplication right? Even a basic matrix multiplication sample could be enough for a sample. The sample could even just display some text. But sobel or any other simple convolution would make it more compelling.
- addedsample requestRequest for a new sampleRequest for a new samplesample wantedWe definitely want to add this sample; contributions welcomeWe definitely want to add this sample; contributions welcome
on Mar 5, 2024 Also it should have a toggle to enable/disable dp4a and hopefully see some performance improvement.
Is dp4a available in the current version of WebGPU?
dp4a is available starting in Chromium M123. So today, that would be Chrome Beta and newer.
If possible, I'd like to try my hand at this issue, at least for the next week (sorry about the timeline, day job is gonna day job). Sobel filter is a good place to start I think.
EDIT: 'GPGPU Compute Category' or 'Features Category'?
Just want to make sure I understand the assignment, the intended use of dp4a here. Instead of writing, say, this for our sobel filter.... (below is in pseudo-wgsl)
(Pixels loaded from texture using textureLoad and global_invocation_id let result = 1 * pixel1.r + 2 * pixel2.r + 1 * pixel 3. r - 1 * textureLoad(inputTexture, vec2<u32>(id.x + 1, id.y - 1), 0).r textureStore(output, id.xy, result)We should do something like this?
let pixelPack = pack4xU8Clamp(pixel1.r, pixel2.r, pixel3.r, pixel4.r) let kernelPack = pack4xU8Clamp(1, 2, 1, -1) let result = dot4U8Packed(pixelPack, kernelPack); textureStore(output, id.xy, result)Also it should have a toggle to enable/disable dp4a and hopefully see some performance improvement.
I suspect a Sobel filter is simple enough that it's limited by memory bandwidth instead of computation.
So I wouldn't get hung up on perf improvement for this sample.Somebody else should take this on, I understand the functionality, but I'm struggling with the quantization of the dp4a result back to something usable.
- added a commit that references this issue
on Jun 24, 2026 - added a commit that references this issue
on Jul 7, 2026 - added a commit that references this issue
on Sep 16, 2026 There's now a very basic example in #572 that lets people play with
dot4I8Packedbut it doesn't do anything that would actually be faster than an unoptimized dot product. Would be interested to see someone either do something more graphical and compute intensive, which could replace this sample.
@beaufortfrancois requested that some small examples be published here showing how to use the new WGSL primitives aimed at AI/ML workloads:
shader-f16, DP4A, and soon, subgroups. Could we consider this?Not sure what would be the most compelling - perhaps something with some visual output, and running a microbenchmark against the fallback WGSL code, assuming the feature is actually supported?