Skip to content

Index the destination once in _copy_subblocks! - #81

Open
lkdvos wants to merge 1 commit into
mainfrom
ld-copy-subblocks
Open

lkdvos wants to merge 1 commit into
mainfrom
ld-copy-subblocks

Conversation

@lkdvos

@lkdvos lkdvos commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

_copy_subblocks! (used when converting a block tensor to a dense TensorMap, e.g. MPSKit's _fuse_env when preparing effective Hamiltonians) looped over every fusion-tree block of the destination and, for each, over every nonzero block of the source, reading it as v[f₁, f₂]. Each of those reads goes through subblockstructure → degeneracystructure, i.e. TensorKit's global LRUCache behind a lock, so the copy did (#fusion trees) × (#nonzero blocks) cache lookups.

This PR

  • iterates the source blocks with subblocks(v) and indexes a single subblocks(tdst) iterator by fusion-tree pair, so each tensor's structure is looked up once;
  • computes the block ranges from per-leg cumulative dimensions (one small vector per leg and sector, computed once), instead of a blockedrange per destination subblock.

Benchmark: copying the left environment × MPO block tensor of an N₂/cc-pVDZ DMRG2 step (D = 300, U₁×SU₂×fℤ₂, 416 nonzero blocks, 931 fusion-tree blocks in the dense result), single-threaded BLAS, results identical:

serial allocated 8 concurrent copies
before 0.537 s 0.107 GiB 2.18 s
after 0.237 s 0.068 GiB 0.386 s

The concurrent case is where the lock contention showed up: with several tasks preparing environments, the old version spent most of its time waiting on the cache's SpinLock.

Also bumps the version to v0.3.20. Full test suite passes locally (Julia 1.11.2).

🤖 Generated with Claude Code

@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

Files with missing lines Coverage Δ
src/tensors/abstractblocktensor/conversion.jl 79.54% <100.00%> (-0.46%) ⬇️

... and 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@lkdvos
lkdvos force-pushed the ld-copy-subblocks branch from a45e397 to b290485 Compare October 4, 2026 01:05
`_copy_subblocks!` looped over the fusion-tree blocks of the destination and read
every nonzero block of the source with `v[f₁, f₂]`, which goes through TensorKit's
lock-protected `degeneracystructure` cache on each call: (#fusion trees) ×
(#nonzero blocks) cache hits. It now iterates the source blocks with `subblocks(v)`
and indexes a single `subblocks(tdst)` iterator by fusion-tree pair, so the
structure is looked up once per tensor. The block ranges come from per-leg
cumulative dimensions computed once, instead of a `blockedrange` per subblock.

Bump version to v0.3.20.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lkdvos
lkdvos force-pushed the ld-copy-subblocks branch from b290485 to 0ab79d4 Compare October 4, 2026 01:23

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant