Hi! We build a voxel sandbox kit on flutter_scene 0.23.0
(voxel_scene /
voxel_game) and spent a few weeks on its frame rate
on a Mac (120 Hz, Retina) and a Galaxy S24 Ultra (Adreno 750, 120 Hz). We measured with
FLUTTER_SCENE_PROFILE, a Dart timeline of profile builds, and the VM's allocation tracer.
Most of what we found was on our side, and we fixed it there. We now share one geometry
per rig part so draws instance, and we merge chunks into 2 × 2 regions (~100 colour draws
became ~40). The terrain uses a packed 16-byte vertex, and fixed groups of boxes are one mesh.
What is left is inside flutter_scene, so here it is, with numbers and suggestions. We are
happy to send PRs for any of these if you tell us which you would take and in what shape.
1. A change to the light direction rebuilds every static shadow cascade in one frame
Our day/night cycle quantises the sun: it moves 2° every 3.33 s (a 600 s day), so that
between steps the static shadow cache holds. Each step still costs one long frame.
DirectionalShadowCache.plan treats any direction change over 1e-10 (squared) as a
parameter change. It clears every entry (render/shadow_cache.dart:101-116), and "a
change to the light basis ... rebuilds the cache outright" (:91-92).
- A fresh entry has no content, so every cascade takes the unamortized
"must render this frame" branch (:132-135).
- Each cascade re-culls the whole static world (
render/shadow_pass.dart:283).
- The new entries have
tile == null, so every cascade also creates a fresh
r32Float texture (render/shadow_pass.dart:246).
Measured
- On the S24 (2 cascades, 1024², 48 m, the whole terrain marked
shadowStatic), every run
shows encode spans of 12.5–22 ms at each sun step, every 3.33 s. The steady encode
is ~4 ms, and none of these frames has a GC or any other work in it (Dart timeline,
profile build).
- At 120 Hz that is 2–3 frames missed every step. These are the phone's slowest frames
that have no garbage collection in them.
Suggestions (any one helps):
- Treat a direction change like a content change. Mark the tiles stale instead of
dropping them, and let the existing amortized path (maxAmortizedRefreshes, nearest
first) refresh them one per frame. plan already returns the effective cascade each tile
was rendered with, so a far cascade would lag the sun by a frame or two.
- Keep the tile textures across a rebuild. Reallocating them buys nothing when the
resolution has not changed.
- Make the thresholds public.
maxAmortizedRefreshes (and maybe an angular tolerance
for the direction) are static const today (:76-79).
The docs recommend cacheStaticShadows = false for a light "whose direction changes every
frame" (light.dart:138-143). A stepped sun is the case in between: static for seconds at
a time, then a small jump.
2. The encode's CPU cost per draw, mostly allocation
On the phone the frame is the UI thread's, and the encode is roughly draws × passes.
flutter_scene's own profile gives these numbers:
- The colour pass cost ~31 µs a draw on the S24 (95–99 draws, 3.7–4.0 ms), against
~4 µs on the Mac.
- Twelve extra draws of one box each cost ~0.3 ms on the phone.
- The cull costs ~2 µs a node on the phone and ~0.4 µs on the Mac.
What the UI isolate allocates while rendering:
- Rates. The VM's allocation tracer, over one still orbiting camera on a loaded world,
measured ~18 MB/s on the S24 and ~21 MB/s on the Mac. With 40 animated creatures
the S24 allocates ~59 MB/s.
- Owners. By first non-SDK library on the stack: flutter_scene 50%, flutter_gpu 23%,
vector_math 16%. By call site, 60% sits under Scene.renderViews. (The tracer
samples, so these shares hold; the rates come from the timeline.)
- Pauses. With the creatures, scavenges run ~4 a second, each up to 2.6 ms, and some
land inside frames.
These are the sites we found, each per draw or per node, every frame:
| Where |
What |
node.dart:1489, :1509-1511 (Node.scenePrePass) |
two capturing closures per node per frame |
scene_encoder.dart:700-707 → sceneSortDepth (:319-328) |
per recorded draw: PerspectiveCamera.forward (camera.dart:317, two Vector3), Aabb3.center (a clone), transformed3 (a copy), worldCenter - cameraPosition (a clone): five Vector3s and their Float32Lists. It is computed even though, for opaque draws, depth is only the last tie-breaker (see 4) |
render/frame_transients.dart:336-373 (TransientArena.emplace) |
per call: the planEmplacement record, two ByteBuffer objects, two Uint8List views, a gpu.BufferView; callers add a ByteData.sublistView (render/instance_packing.dart:771, per non-instanced draw) |
flutter_gpu Shader.getUniformSlot |
a new UniformSlot per call, not cached; called per draw by the PBR bind (physically_based_material.dart:1620, :1672, :1760) and by any custom Geometry.bind |
scene_encoder.dart:355 (resolvePipeline) |
a record key per draw |
shader_uniform_bindings.dart:25, :55-66; sky_sources.dart:78 |
Float32List.fromList of a spread list per bind; map-entry iteration and a default SamplerOptions per bind |
The biggest single item for a custom Geometry is the per-draw FrameInfo. The
encoder memoizes it per shader and depth bias, but only for UnskinnedGeometry
(scene_encoder.dart:789-804). Any other Geometry gets a full bind every draw
(:805-814), so it emplaces the same 80-byte FrameInfo once per draw, with a fresh
UniformSlot each time.
In our trace, that one site was 13% of the UI isolate's allocations. We cannot cache
it ourselves: TransientWriter exposes no frame boundary, and arena blocks can recycle
mid-frame (frame_transients.dart:404-432).
Suggestions:
- Split
Geometry.bind into bindBuffers and bindFrameInfo, or give the memo a hook,
so a custom geometry gets the same skip UnskinnedGeometry has.
- Cache
UniformSlots per shader and name (in flutter_gpu, or a small map in flutter_scene).
- Reuse scratch vectors in
_depthOf, and compute an item's sort depth only when the sort
needs it.
- Make
emplace write through one cached Uint8List per block.
- Walk children and components in
scenePrePass with a loop instead of capturing closures.
3. A supported path for a custom Geometry (and fewer bytes for MeshGeometry)
MeshGeometry.fromArrays always uploads six streams, 72 bytes a vertex, filling the
absent ones with defaults (geometry/geometry.dart:1059-1081,
geometry/interleaved_layout.dart:365-380), and keeps a CPU copy by default. Our
terrain uses position, normal, colour and two light values.
We moved it to a custom Geometry with two uint32x2 streams (16 bytes) and its own
vertex and depth-only shaders:
- RSS −15% at a 12-chunk radius on the Mac (503 → 426 MB), and −28 to −43 MB on the
phone at radius 6.
- +11% fps where the Mac draws 3.7× the faces (84 → 94 fps).
- The Mac's GPU time at
dpr 2 did not move (it pays for pixels there).
It works, but only through internals:
setVertexStreams and bindGeometryBuffers are @internal.
Geometry.bind is typed with the src/gpu/gpu.dart shim, so we import src/ and pin
flutter_scene to an exact version.
Suggestions:
- Make the custom-
Geometry contract public: stream setup, buffer binding, depthOnlyVertex
and the standard varyings a vertex shader must write.
- Optionally, have
MeshGeometry skip the streams it was not given, with a layout per
combination.
4. Opaque draws are not front to back across geometries (a question)
The class doc says opaque draws are "sorted by pipeline ... and then front-to-back"
(scene_encoder.dart:388-390). The comparator's keys are pipeline, material, geometry
(identityHashCode), light list, channels and fade, and only then depth
(scene_encoder.dart:1058-1080).
So between different geometries (every chunk region of a world, every distinct prop), the
order follows identityHashCode, not distance.
We have not measured what front-to-back would buy. On Apple's GPUs (hidden-surface
removal) and Adreno (LRZ) we expect little. On an immediate-mode desktop GPU it is the
classic early-Z win.
The question: is depth-before-geometry intended to be traded for batching? One option
is to sort by pipeline and material, then by coarse depth buckets, then by geometry. That
keeps batching and gets most of the early-Z benefit. If you only want the doc fixed, that
is fine too.
Not asked: pipeline warm-up
The first run after an install stops the UI thread for ~0.6 s (S24) to ~0.8 s (Mac), in
two encodes that compile pipelines. We had this down as a request, until we found
Scene.warmUp(views, includeOffscreen: true) (scene.dart:1397). That one is on us.
Thanks for flutter_scene. The numbers above come from A/B runs we can share (JSONL, per
run) if they help.
Hi! We build a voxel sandbox kit on flutter_scene 0.23.0
(voxel_scene /
voxel_game) and spent a few weeks on its frame rate
on a Mac (120 Hz, Retina) and a Galaxy S24 Ultra (Adreno 750, 120 Hz). We measured with
FLUTTER_SCENE_PROFILE, a Dart timeline of profile builds, and the VM's allocation tracer.Most of what we found was on our side, and we fixed it there. We now share one geometry
per rig part so draws instance, and we merge chunks into 2 × 2 regions (~100 colour draws
became ~40). The terrain uses a packed 16-byte vertex, and fixed groups of boxes are one mesh.
What is left is inside flutter_scene, so here it is, with numbers and suggestions. We are
happy to send PRs for any of these if you tell us which you would take and in what shape.
1. A change to the light direction rebuilds every static shadow cascade in one frame
Our day/night cycle quantises the sun: it moves 2° every 3.33 s (a 600 s day), so that
between steps the static shadow cache holds. Each step still costs one long frame.
DirectionalShadowCache.plantreats any direction change over1e-10(squared) as aparameter change. It clears every entry (
render/shadow_cache.dart:101-116), and "achange to the light basis ... rebuilds the cache outright" (
:91-92)."must render this frame" branch (
:132-135).render/shadow_pass.dart:283).tile == null, so every cascade also creates a freshr32Floattexture (render/shadow_pass.dart:246).Measured
shadowStatic), every runshows
encodespans of 12.5–22 ms at each sun step, every 3.33 s. The steady encodeis ~4 ms, and none of these frames has a GC or any other work in it (Dart timeline,
profile build).
that have no garbage collection in them.
Suggestions (any one helps):
dropping them, and let the existing amortized path (
maxAmortizedRefreshes, nearestfirst) refresh them one per frame.
planalready returns the effective cascade each tilewas rendered with, so a far cascade would lag the sun by a frame or two.
resolution has not changed.
maxAmortizedRefreshes(and maybe an angular tolerancefor the direction) are
static consttoday (:76-79).The docs recommend
cacheStaticShadows = falsefor a light "whose direction changes everyframe" (
light.dart:138-143). A stepped sun is the case in between: static for seconds ata time, then a small jump.
2. The encode's CPU cost per draw, mostly allocation
On the phone the frame is the UI thread's, and the encode is roughly draws × passes.
flutter_scene's own profile gives these numbers:
~4 µs on the Mac.
What the UI isolate allocates while rendering:
measured ~18 MB/s on the S24 and ~21 MB/s on the Mac. With 40 animated creatures
the S24 allocates ~59 MB/s.
vector_math 16%. By call site, 60% sits under
Scene.renderViews. (The tracersamples, so these shares hold; the rates come from the timeline.)
land inside frames.
These are the sites we found, each per draw or per node, every frame:
node.dart:1489,:1509-1511(Node.scenePrePass)scene_encoder.dart:700-707→sceneSortDepth(:319-328)PerspectiveCamera.forward(camera.dart:317, twoVector3),Aabb3.center(a clone),transformed3(a copy),worldCenter - cameraPosition(a clone): fiveVector3s and theirFloat32Lists. It is computed even though, for opaque draws, depth is only the last tie-breaker (see 4)render/frame_transients.dart:336-373(TransientArena.emplace)planEmplacementrecord, twoByteBufferobjects, twoUint8Listviews, agpu.BufferView; callers add aByteData.sublistView(render/instance_packing.dart:771, per non-instanced draw)Shader.getUniformSlotUniformSlotper call, not cached; called per draw by the PBR bind (physically_based_material.dart:1620,:1672,:1760) and by any customGeometry.bindscene_encoder.dart:355(resolvePipeline)shader_uniform_bindings.dart:25,:55-66;sky_sources.dart:78Float32List.fromListof a spread list per bind; map-entry iteration and a defaultSamplerOptionsper bindThe biggest single item for a custom
Geometryis the per-drawFrameInfo. Theencoder memoizes it per shader and depth bias, but only for
UnskinnedGeometry(
scene_encoder.dart:789-804). Any otherGeometrygets a fullbindevery draw(
:805-814), so it emplaces the same 80-byteFrameInfoonce per draw, with a freshUniformSloteach time.In our trace, that one site was 13% of the UI isolate's allocations. We cannot cache
it ourselves:
TransientWriterexposes no frame boundary, and arena blocks can recyclemid-frame (
frame_transients.dart:404-432).Suggestions:
Geometry.bindintobindBuffersandbindFrameInfo, or give the memo a hook,so a custom geometry gets the same skip
UnskinnedGeometryhas.UniformSlots per shader and name (in flutter_gpu, or a small map in flutter_scene)._depthOf, and compute an item's sort depth only when the sortneeds it.
emplacewrite through one cachedUint8Listper block.scenePrePasswith a loop instead of capturing closures.3. A supported path for a custom
Geometry(and fewer bytes forMeshGeometry)MeshGeometry.fromArraysalways uploads six streams, 72 bytes a vertex, filling theabsent ones with defaults (
geometry/geometry.dart:1059-1081,geometry/interleaved_layout.dart:365-380), and keeps a CPU copy by default. Ourterrain uses position, normal, colour and two light values.
We moved it to a custom
Geometrywith twouint32x2streams (16 bytes) and its ownvertex and depth-only shaders:
phone at radius 6.
dpr2 did not move (it pays for pixels there).It works, but only through internals:
setVertexStreamsandbindGeometryBuffersare@internal.Geometry.bindis typed with thesrc/gpu/gpu.dartshim, so we importsrc/and pinflutter_scene to an exact version.
Suggestions:
Geometrycontract public: stream setup, buffer binding,depthOnlyVertexand the standard varyings a vertex shader must write.
MeshGeometryskip the streams it was not given, with a layout percombination.
4. Opaque draws are not front to back across geometries (a question)
The class doc says opaque draws are "sorted by pipeline ... and then front-to-back"
(
scene_encoder.dart:388-390). The comparator's keys are pipeline, material, geometry(
identityHashCode), light list, channels and fade, and only then depth(
scene_encoder.dart:1058-1080).So between different geometries (every chunk region of a world, every distinct prop), the
order follows
identityHashCode, not distance.We have not measured what front-to-back would buy. On Apple's GPUs (hidden-surface
removal) and Adreno (LRZ) we expect little. On an immediate-mode desktop GPU it is the
classic early-Z win.
The question: is depth-before-geometry intended to be traded for batching? One option
is to sort by pipeline and material, then by coarse depth buckets, then by geometry. That
keeps batching and gets most of the early-Z benefit. If you only want the doc fixed, that
is fine too.
Not asked: pipeline warm-up
The first run after an install stops the UI thread for ~0.6 s (S24) to ~0.8 s (Mac), in
two encodes that compile pipelines. We had this down as a request, until we found
Scene.warmUp(views, includeOffscreen: true)(scene.dart:1397). That one is on us.Thanks for flutter_scene. The numbers above come from A/B runs we can share (JSONL, per
run) if they help.