Skip to content

[Refactor] Remove dead runtime-device memory wrappers - #381

Closed
Cstandardlib wants to merge 49 commits into
abacusmodeling:developfrom
deepmodeling:fix/remove-dead-device-memory-wrappers
Closed

Cstandardlib wants to merge 49 commits into
abacusmodeling:developfrom
deepmodeling:fix/remove-dead-device-memory-wrappers

Conversation

@Cstandardlib

Copy link
Copy Markdown
Contributor

Linked Issue

Fix #8039

What's changed?

Delete the five AbacusDevice_t-keyed wrappers (resize_memory, set_memory, synchronize_memory, cast_memory, delete_memory). They have no call site in the tree or anywhere in the history, and their runtime device dispatch duplicated the compile-time *_op templates, which is why the defective `||` conditions behind #7553 went unnoticed.

Also drop #7941's explicit instantiations, the then-unused tool_quit.h include, and the dispatch tests that exercised the removed API. The *_op struct templates are untouched.

mohanchen and others added 30 commits September 16, 2026 09:22
…rce/stress free-function extraction, and nonlocal/InfoNonlocal modernization (#7961)

* refactor(hamilt): remove GlobalV/PARAM deps from HamiltLCAO

Snapshot inp.nspin/inp.vl_in_h as members at construction so
getHR_vector/updateHk/refresh no longer read global PARAM, and pass the
EXX restart flag into the constructor as load_exx_flag instead of
reading GlobalC::restart. Drop the now-unused global_variable.h include.
Call sites in esolver_ks_lcao/esolver_double_xc/lcao_others compute the
flag with the original logic, so behavior is unchanged.

Verified: make -j 30 in build_max_para_test builds with zero errors.

* refactor(hamilt): drop unused default arg in updateSk

All call sites pass hk_type explicitly, so the default argument was dead
weight. Removing it aligns with the no-default-arguments rule without
changing any caller or behavior.

Verified: make -j 30 in build_max_para_test builds with zero errors.

* refactor(hamilt): cache OperatorLCAO downcast in HamiltLCAO

updateHk and refresh repeated the same dynamic_cast<OperatorLCAO*> four
times; cache the result in a private ops_lcao_ member filled on first
use via getOperatorLCAO(). Also spell out the matrix() local type
explicitly instead of auto.

Verified: make -j 30 in build_max_para_test builds with zero errors.

* refactor(hamilt): split HamiltLCAO constructor operator-chain branches

Move the gamma-only and multi-k operator-chain construction out of the
~300-line constructor into private init_gamma_operators() /
init_multik_operators(), cutting constructor cyclomatic complexity from
35 to 12. The TDDFT velocity-gauge block (TDEkinetic/TDNonlocal) stays
guarded by std::is_same<TK, complex> inside init_multik_operators so the
double instantiation dead-branch-eliminates it -- those operators have
no double instantiation, and hoisting them into a standalone function
would produce undefined references.

Verified: make -j 30 in build_max_para_test builds and links with zero
errors.

* refactor(hamilt): dedupe DFTU/DeePKS operator construction

The DFT+U and DeePKS operator blocks were byte-identical in the gamma
and multi-k branches. Extract them into private add_dftu_op() and
add_deepks_op() helpers so each is defined once, removing ~60 lines of
duplication and slightly lowering init_multik_operators complexity.
Both operators have double and complex instantiations, so the helpers
are safe for every HamiltLCAO specialization.

Verified: make -j 30 in build_max_para_test builds and links with zero
errors.

* refactor(hamilt): tidy HamiltLCAO misc cleanups

- drop trailing return; at end of the constructor
- make the TDDFT nonlocal term conditional up front instead of
  new-then-maybe-delete, removing a new/delete pair
- remove trailing spaces on the include guard
- drop redundant virtual on updateHk (override already implies it)
- unify destructor to plain delete ops/hR/sR (delete nullptr is safe)
- remove dead member const int istep = 0 (never read; the ctor
  parameter shadows it and is forwarded to OperatorEXX)

All changes are behavior-preserving cleanups; no allocation ownership
path changes, so no memory leaks introduced.

Verified: make -j 30 in build_max_para_test builds and links with zero
errors.

* refactor(hamilt): prune unused includes and order by call chain

Drop six unused headers verified by grep + full build:
- source_io/module_parameter/parameter.h (no PARAM usage)
- source_hamilt/module_xc/xc_functional.h (no XC_Functional usage)
- source_hsolver/hsolver_lcao.h and diago_elpa.h (no solver symbols)
- module_operator_lcao/meta_lcao.h (no Meta node constructed)
- module_operator_lcao/op_exx_lcao.h duplicate (already under __EXX)

Keep dspin_lcao.h: DeltaSpin is declared there (class name does not
match the file name, so it survived an initial over-pruning caught by
the build). Reorder the remaining includes to follow the constructor
call chain: infra -> dftu base/setup -> electronic state -> operators
(overlap -> kinetic -> nonlocal -> veff -> dftu -> dspin -> tddft).

Verified: make -j 30 in build_max_para_test builds and links with zero
errors.

* refactor(hamilt): extract LCAO operator-chain construction to factory

HamiltLCAO carried a construction-time factory responsibility (building
the overlap/kinetic/nonlocal/veff/DFTU/DeePKS/TDDFT/spin-constrain
operator chain) that is independent of the object's runtime state. Move
that logic out of the class into this-free factory functions, matching
the explicit-parameter style used elsewhere (KListIO, dftu_pw).

New hamilt_lcao_factory.{h,cpp} in namespace hamilt:
- LcaoOpsBundle<TK,TR>: the only two construction products -- the chain
  head (ops) and the DeePKS V_delta(R) handle. hR/sR/hsk stay allocated
  by the caller and are passed in as inputs, keeping ownership clear.
- build_gamma_ops / build_multik_ops: free functions with explicit
  parameters; add_dftu_op / add_deepks_op move to an anonymous namespace
  as internal helpers that append onto the chain head by reference.

HamiltLCAO constructor now calls the factory and assigns bundle.ops /
bundle.v_delta_R; the four private builder method declarations are
removed from the header, and the now-unused operator-node includes are
pruned from hamilt_lcao.cpp.

Explicit instantiation (3 TK/TR combos x 2 functions) keeps the
complex-only TDEkinetic/TDNonlocal guard inside build_multik_ops so the
double instantiation still dead-branch-eliminates those references.

Verified: make -j 30 in build_max_para_test builds and links with zero
errors after both the split and the include pruning.

* refactor(hamilt): guard dft_plus_u with explicit branches and WARNING_QUIT

add_dftu_op previously used if (==2) ... else ..., which silently routed
every non-2 value (including invalid ones) into the first-zeta NAO DFTU
implementation. Per input semantics (dft_plus_u: 0 = off, 1 =
radius-adjustable, 2 = first-zeta NAO), make the branches explicit:

  == 1 -> DFTU (radius-adjustable, default new method, listed first)
  == 2 -> OperatorDFTU (first-zeta NAO, old method kept for testing)
  else -> ModuleBase::WARNING_QUIT on any out-of-range value

add_dftu_op is only reachable when dft_plus_u != 0, so the else branch
turns previously-silent misclassification into a clear abort. Adds the
source_base/global_function.h include for ModuleBase::WARNING_QUIT.

Verified: make -j 30 in build_max_para_test builds and links with zero
errors; valid inputs (1/2) keep identical behavior.

* refactor(hamilt): manage HamiltLCAO::hsk with std::unique_ptr

hsk is exclusively owned by HamiltLCAO (allocated in the SCF
constructor, freed in the destructor, only read elsewhere), so hold it
in a std::unique_ptr instead of a raw pointer. Pass .get() to the
operator-chain constructors and factories, and drop the manual delete.

C++11 baseline: use reset(new ...) instead of std::make_unique.

Verified: builds with make -j 30.

* refactor(hamilt): manage HamiltLCAO hR/sR with std::unique_ptr

hR and sR are exclusively owned by HamiltLCAO (allocated in the
constructors, freed in the destructor, only read elsewhere; no caller
rebinds or deletes them). Hold them in std::unique_ptr and drop the
manual deletes.

This requires getHR()/getSR() to return HContainer<TR>* by value
instead of HContainer<TR>*&, since a unique_ptr member cannot expose a
reference to its stored pointer. No call site relies on the reference
(verified: nothing assigns to or rebinds through getHR()/getSR()), so
the change is behavior-compatible.

C++11 baseline: use reset(new ...) instead of std::make_unique.

Verified: builds with make -j 30.

* fix makefile

* remove useless TAC

* refactor(dftu): rename DFTU operator classes for clarity

Rename DFTU<OperatorLCAO<TK,TR>> to DFTU_onsite and
OperatorDFTU<OperatorLCAO<TK,TR>> to DFTU_firstzeta to better
reflect the two DFT+U projection methods (radius-adjustable on-site
vs first-zeta NAO) and to match the naming style of other LCAO
operators (Overlap, Nonlocal, etc.).

* refactor(dftu): replace GlobalFunc::ZEROS with std::fill

Remove the ModuleBase::GlobalFunc::ZEROS dependency in module_dftu by
using std::fill on the raw buffers, consistent with the preference for
std::fill/std::copy over ZEROS/COPYARRAY. T(0) covers both the double
and std::complex<double> instantiations of cal_pot_onsite/cal_pot_uterm.

* refactor(dftu): drop __DEBUG guard around input asserts

Keep the nspin/null-pointer and nlm-size asserts active in all builds;
they validate cheap invariants, not expensive debug-only checks.

* fix(hamilt): allocate hR in HamiltLCAO vacuum constructor

The vacuum constructor documented "only HR and SR will be initialed as
empty HContainer" but only allocated sR, passing an unallocated hR to
the Overlap node. With raw pointers this was an uninitialized-value UB;
with unique_ptr it is a null pointer that would segfault if any caller
invokes init(). Allocate hR alongside sR to match the documented
contract. hsk stays null because the vacuum path has no k-space matrix
and never calls init().

No numerical change: the sole caller (esolver_gets) only invokes
contributeHR(), which writes SR and never reads hR/hsk.

* refactor(lcao): replace ForceStressArrays raw pointers with std::vector (steps 1-3)

Step 1: eliminate DSloc_R* aliasing in cal_dS by writing DHloc_fixedR_*
directly in single_derivative ('S' branch), guarded by write_dsloc_r.

Step 2: convert 12 gamma-only stress members (DSloc_11..33,
DHloc_fixed_11..33) from double* to std::vector<double>.

Step 3: convert 7 multi-k stress members (DH_r, stvnl11..33) from
double* to std::vector<double>; replace OpenMP ZEROS lambda with
resize(n, 0.0); update nullptr checks to .empty().

* refactor(lcao): convert DSloc_x/y/z and DHloc_fixed_x/y/z to std::vector (step 4)

Replace 6 gamma-only force members from double* to std::vector<double>.
Update all call sites to use .data() for set_force and cal_pulay_fs,
and nullptr checks to .empty() in check_folded_arrays.

* refactor(lcao): convert DHloc_fixedR_x/y/z to std::vector (step 5)

Replace 3 multi-k force members from double* to std::vector<double>.
Resize with zero-init replaces new + ZEROS + OpenMP lambda in
force_lcao_k.cpp and spar_dh.cpp. Remove all corresponding delete[].
Build verified in build_max_para_test.

* refactor(lcao): convert DSloc_Rx/Ry/Rz to std::vector (step 6)

Replace 3 multi-k force members from double* to std::vector<double>.
Update write_dsloc_r guard from nullptr to .empty() in
single_derivative. Update check_folded_arrays nullptr checks.
Build verified in build_max_para_test.

* refactor(lcao): replace InfoNonlocal raw pointers with std::vector

Convert all raw new/delete arrays in InfoNonlocal and its local
temporaries to std::vector, eliminating manual memory management:

- InfoNonlocal::Beta: Numerical_Nonlocal* -> std::vector<Numerical_Nonlocal>
- InfoNonlocal::nproj: int* -> std::vector<int>
- Set_NonLocal/Read_NonLocal local arrays -> std::vector
- setupNonlocal(): use resize/assign instead of delete[]+new[]
- Update all call sites to use .data() for raw-pointer interfaces
- Align setup_nonlocal.h style (4-space indent, unified comments)

Files updated:
- source/source_lcao/setup_nonlocal.h/.cpp
- source/source_lcao/lcao_init_basis.cpp
- source/source_esolver/esolver_lr_lcao_tddft.cpp
- source/source_lcao/module_operator_lcao/test/test_{t_nl_cd,nonlocal}.cpp
- source/source_lcao/module_rt/test/snap_psb_half_tddft_test.cpp
- source/source_lcao/module_operator_lcao/test/tmp_mocks.cpp

* refactor(lcao): eliminate GlobalV/PARAM dependencies in InfoNonlocal

Pass my_rank, log stream, and out_element_info as explicit parameters
instead of reading GlobalV::MY_RANK, GlobalV::ofs_running, and
PARAM.inp.out_element_info directly. This aligns with ABACUS governance
rule 1 (no cross-layer control through globals).

Changes:
- Set_NonLocal: add my_rank parameter, use it for plot() calls
- Read_NonLocal: add out_element_info and log parameters
- setupNonlocal: add my_rank parameter, forward to callees
- Remove parameter.h include (no longer needed)
- Update LCAONonlocalInfo::setupNonlocal wrapper signature
- Update all call sites to pass GlobalV::MY_RANK explicitly
- Update snap_psb_half_tddft_test.cpp Set_NonLocal calls with my_rank=0

Verified: make -j 30 in build_std_gpu passes.
Quality score: setup_nonlocal.cpp 54 -> 68 (global_dependency eliminated).

* refactor(lcao): extract functions to reduce cyclomatic complexity

Extract 5 helper functions from Set_NonLocal and Read_NonLocal:

- build_soc_coefficients: SOC coefficient matrix construction (from Set_NonLocal)
- build_beta_r: radial projector truncation and copy (from Set_NonLocal)
- read_header: parse <HEADER> section (from Read_NonLocal)
- read_dij: parse <DIJ> section (from Read_NonLocal)
- read_projector: parse one <PP_BETA> block (from Read_NonLocal)

Also remove dead code: coefficient_D_in and coefficient_D_nc_in
in Read_NonLocal were written but never read.

Cyclomatic complexity:
- Set_NonLocal: 19 -> eliminated (main body now <10)
- Read_NonLocal: 24 -> eliminated (main body now <10)
- build_soc_coefficients: 14 (extracted, can be further split)

Verified: make -j 30 in build_std_gpu passes.
Quality score: setup_nonlocal.cpp 68 -> 85.

* refactor(lcao): encapsulate InfoNonlocal member variables

Convert 4 public member variables to private and add const getters/setters:
- Beta -> get_Beta(), get_Beta(it), get_Beta_data(), resize_Beta()
- nproj -> get_nproj(), get_nproj(it), assign_nproj()
- nprojmax -> get_nprojmax(), set_nprojmax()
- rcutmax_Beta -> get_rcutmax_Beta(), set_rcutmax_Beta()

Update all external call sites to use getters/setters instead of
direct member access. LCAONonlocalInfo now uses the new interface.

Verified: make -j 30 in build_std_gpu passes.
Quality score: setup_nonlocal.h 85 -> 87, lcao_nonlocal_info.h 96.

* refactor(lcao): remove Read_NonLocal dead code and helpers

Remove Read_NonLocal, read_header, read_dij, and read_projector
which were unreachable because readin_nonlocal was hardcoded to false.
This eliminates ~300 lines of dead code including all NONLOCAL file
parsing logic.

Also remove the readin_nonlocal branch from setupNonlocal, keeping
only the Set_NonLocal path.

Verified: make -j 30 in build_std_gpu passes.
Quality score: setup_nonlocal.cpp 85 -> 90, setup_nonlocal.h 87 -> 92.

* fix(lcao): add get_nproj_ref for non-const lvalue reference

Set_NonLocal takes int& n_projectors which requires a modifiable
lvalue. get_nproj(it) returns int by value which cannot bind.
Add get_nproj_ref(it) that returns int& for this use case.

Update snap_psb_half_tddft_test.cpp to use get_nproj_ref(0) in
both Set_NonLocal call sites.

Verified: make -j 30 in build_std_gpu passes.

* fix(tddft): construct Nonlocal for hR pair insertion in velocity gauge

Nonlocal::initialize_HR inserts atom pairs into hR using a cutoff that
includes the nonlocal pseudopotential radius, which may be larger than
the orbital cutoff used by EKinetic/Veff. TDEkinetic and TDNonlocal both
build hR_tmp by iterating over hR's pairs, so skipping Nonlocal's
construction in TDDFT velocity gauge mode left hR with missing pairs,
producing an incomplete hR_tmp and incorrect Hamiltonian.

Restore the original pattern: always construct Nonlocal when vnl_in_h
is set, then conditionally add it to the operator chain (or delete it).

* fix(lcao): forbid copying Numerical_Nonlocal and avoid vector reallocation

InfoNonlocal::Beta was changed from a raw array to
std::vector<Numerical_Nonlocal> in the recent refactor. Since
Numerical_Nonlocal owns a raw Proj buffer but defines no copy
semantics, Beta.resize() reallocating would shallow-copy elements,
leaving dangling Proj pointers and causing SEGFAULTs in all
module_deepks unit tests (and any run with ntype > 1).

Fix without introducing copy/move semantics:
- Explicitly delete Numerical_Nonlocal copy constructor and copy
  assignment, so any accidental copy now fails at compile time.
- Replace Beta wholesale via move-assigning a fresh vector instead of
  resize(), so elements are constructed in place and never relocated.
- Drop the Beta.resize(1) preallocation in InfoNonlocal's constructor
  (also in the operator_lcao test mock) to keep the invariant.

Verification: static analysis only; build and test run not performed.

* fix(test): allocate Beta before direct Set_NonLocal calls in tddft test

snap_psb_half_tddft_test calls InfoNonlocal::Set_NonLocal directly
without setupNonlocal(), which is the only path that used to size the
Beta array. After the raw array was replaced by std::vector and the
resize(1) preallocation was removed, Beta was empty and Beta[it] was
out of bounds, causing SEGFAULT in MODULE_LCAO_tddft_snap_psibeta_half_test.

Add resize_Beta(1) next to the existing assign_nproj(1, 0) in both
fixture SetUp() functions.

Verification: static analysis only; build and test run not performed.

* refactor(lcao): split Record_adj::for_2d and deduplicate adjacency check

Extract the copy-pasted direct-cutoff / beta-bridge adjacency test into a
single file-local is_adjacent helper shared by both passes, and split the
~230-line for_2d (cyclomatic complexity 31) into count_adjacent,
allocate_info, and fill_info orchestrated by a thin for_2d wrapper.
Public members and the int*** info layout are unchanged so downstream
consumers need no modifications.

Code quality score for record_adj.cpp: 59 -> 84.

* refactor(lcao): pass npol explicitly to Record_adj::for_2d

Remove Record_adj's reads of PARAM.globalv.npol, PARAM.inp.out_level and
GlobalV::ofs_running. npol is now an explicit argument of for_2d /
count_adjacent (no default argument per governance), and the ParaV.nnr
log is emitted by the three callers instead. Add a public const
getAdjacentInfo() observer on Grid_Driver to expose adj_info read-only.

Code quality score for record_adj.cpp: 84 -> 96 (global_dependency gone).

* refactor(lcao): flatten Record_adj info storage into a single vector

Replace the manually managed int*** info (and the raw int* na_each /
iat2ca) with containers. Adjacent records are stored flat in one
std::vector<std::array<int,5>> with an info_offset prefix-sum table,
exposed read-only through get_info(iat, cb). This removes all raw
new/delete, the info_modified flag, and the nested-vector pointer
chasing, and keeps each atom's records contiguous for the OpenMP fill
loop. na_each / iat2ca become std::vector<int>. Update the three
consumers (density_matrix_io, td_current_io, pulay_fs_temp) and the
dm_r_init test to the new layout.

Code quality score for record_adj.cpp: 96 -> 100 (raw_new_keyword gone).

* refactor(lcao): use injected inp_->out_level at for_2d call sites

The ParaV.nnr log moved out of Record_adj in the previous commit read
PARAM.inp.out_level at each caller, which raised the PR-level global
dependency budget. All three callers already hold an injected INPUT
pointer (this->inp_), so read out_level from it instead of PARAM.inp,
removing three PARAM.inp references.

* refactor(lcao): remove dead Force_Stress_LCAO::integral_part

The two integral_part specializations were the only callers of
Force_LCAO<T>::ftable, and integral_part itself has no callers since
the operator-based force/stress path took over. Remove the dead entry
point first so the ftable implementations can be deleted next.

Verified: make -j 30 in build_max_para_test passes (100% Built target
abacus_max_para).

* refactor(lcao): delete dead force_lcao_gamma.cpp and force_lcao_k.cpp

After removing Force_Stress_LCAO::integral_part (the only caller of
Force_LCAO<T>::ftable), the allocate/ftable/finish_ftable
specializations in these two files have no remaining callers. Delete
the files and drop them from CMakeLists.txt and Makefile.Objects.
The active force/stress path uses operator-based cal_force_stress
plus PulayForceStress::cal_pulay_fs directly.

Verified: make -j 30 in build_max_para_test passes (100% Built target
abacus_max_para).

* refactor(lcao): drop dead Force_LCAO method declarations

With ftable/allocate/finish_ftable deleted, their declarations plus
the never-defined average_force/cal_fedm/cal_ftvnl_dphi/cal_fvl_dphi
declarations are dead. Force_LCAO now only carries the actively used
cal_edm and its ParaV/pot members. Remove the declarations and the
includes (matrix.h, two_center_bundle.h, force_stress_arrays.h,
setup_deepks.h) that only served them.

Verified: make -j 30 in build_max_para_test passes (100% Built target
abacus_max_para).

* refactor(lcao): drop dead DSloc_*/DHloc_fixed_* stress arrays

The DSloc_11/12/13/22/23/33 and DHloc_fixed_11/12/13/22/23/33 arrays
were only written by the gamma-only cal_stress branch of
single_derivative via set_stress, and never read anywhere. The active
LCAO stress path computes the overlap/kinetic/nonlocal contribution
through the operator-based cal_force_stress instead. Remove the 12
arrays from ForceStressArrays, the set_stress call site, and the now
unused set_stress declaration/implementation. single_derivative keeps
its cal_stress parameter because the multi-k branch still uses it to
fill DH_r and stvnl*.

Verified: make -j 30 in build_max_para_test passes (100% Built target
abacus_max_para).

* fix(lcao): add missing TwoCenterBundle include for CUDA build

Forward-declare TwoCenterBundle in force_stress_lcao.h and explicitly
include two_center_bundle.h in force_stress_lcao.cpp to fix CUDA
compilation where the indirect include chain is broken.

* refactor(lcao): rename misplaced .hpp headers to .h

lcao_hs_arrays.hpp is a pure declaration header and the two
pulay_fs_*.hpp files hold template implementations; none of them are
.hpp implementation headers in the prohibited sense. Rename them to .h
and update the 11 include sites so the hpp_implementation rule no
longer flags them.

Quality score: lcao_hs_arrays 34->84, pulay_fs_temp 23->73,
pulay_fs_gint 46->96.

* refactor(lcao): pass gamma_only_local/nspin/npol into sparse_format

sparse_format::cal_dH/cal_dS/cal_dSTN_R/destroy_dH_R_sparse read
PARAM.globalv.gamma_only_local, PARAM.inp.nspin and PARAM.globalv.npol
directly. Pass them as explicit arguments so the functions no longer
depend on global INPUT state, in line with the rule that cross-layer
control through PARAM should not grow.

The remaining PARAM.globalv.nlocal in cal_dH is kept because the
caller has no local value for it; threading it further would only move
the global read, not remove it.

Quality score: spar_dh.cpp 58->77.

* refactor(lcao): pack single_overlap/single_derivative args into ST_env/ST_elem

single_overlap and single_derivative each took 29 parameters, mixing
three kinds of state: the read-only build environment (basis, parallel
layout, unit cell, spin config), the per-element inputs (operator type,
orbital and angular-momentum indices, displacement) and the outputs.

Pack them into two aggregate types in LCAO_domain:
- ST_env: everything fixed for one build_ST_new call, built once before
  the omp region. This also removes the PARAM.globalv.gamma_only_local
  reads inside both functions.
- ST_elem: the per-matrix-element inputs, built once per inner-loop
  iteration and shared by both call sites.

Dead parameters tau1/tau2 (only dtau was used) are dropped. Local index
variables are lowercased (t1/l1/n1/i1, mm1/mm2 for the magnetic quantum
number to avoid clashing with the m1/m2 indices).

The functions stay in lcao_set_st.cpp so they remain inlinable at their
hot inner-loop call sites; the parameter unpacking is POD and optimises
away. Quality score: lcao_domain.h 36->82.

* refactor(lcao): split single_deriv S/T branches into helpers

single_deriv (renamed from single_derivative) had cyclomatic complexity
24, all of it in the multi-k branch that dispatches on operator type (S/T)
x nspin (1/2/4) x spin index is. Extract the per-element writes into two
static free functions, set_deriv_s and set_deriv_t, kept in this
translation unit so they stay inlinable at the hot inner-loop call site.
The main function now only computes the spin index and dispatches.

The nspin==4 S-branch "write olm or write zero" blocks were two symmetric
if/else arms; collapse them to is==0 ? olm[i] : 0.0.

Also fix WARNING_QUIT labels that named LCAO_domain::build_ST_new from
inside set_deriv_s/set_deriv_t/single_overlap/single_deriv; they now name
the function actually raising them so the log points at the right place.

Quality score: lcao_set_st.cpp 45 -> 59.

* refactor(lcao): extract per-element nonlocal accumulators

Split the energy/force accumulation of one <psi|beta><beta|psi> matrix
element out of build_Nonlocal_mu_new into three static helpers
(accum_nlm_energy/accum_nlm_force_soc/accum_nlm_force). Bundle the
call-invariant inputs into NL_env and the per-element indices into
NL_elem so the helpers take explicit arguments instead of reaching into
the enclosing loop, and drop the four nlm_cur*_e/f pointer aliases.
Lowers the function cyclomatic complexity from 60 to 40.

* refactor(lcao): extract build_psi_beta from build_Nonlocal_mu_new

Move the <psi|beta> (and <d psi|beta>) generation loop into a static
build_psi_beta helper, keeping its OpenMP-parallel iat loop inside the
helper and passing nlm_tot/nlm_tot1 out by reference. The main function
now only drives Step 2, lowering build_Nonlocal_mu_new cyclomatic
complexity from 40 to 29.

* refactor(lcao): extract Step2 inner orbital loop into accum_nlm_block

Move the (j, k) orbital loop that dispatches to the energy/force
accumulators out of build_Nonlocal_mu_new into a static accum_nlm_block
helper, packing the per-neighbour inputs (atoms, orbital offsets, iat
slot and the two <psi|beta> block keys) into a file-local NL_pair struct.
Also normalise the remaining K&R brace placement in this file. Lowers
build_Nonlocal_mu_new cyclomatic complexity from 29 to 21.

* refactor(lcao): extract force/stress assembly into helpers

Move the 32 local force/stress part matrices in getForceStress into
LCAOForceParts/LCAOStressParts containers, and extract the force
assembly+print and stress assembly+print blocks into
assemble_and_print_force / assemble_and_print_stress.

* refactor(lcao): extract per-term force/stress calculators from getForceStress

Split the body of getForceStress into focused helpers: cal_operator_fs
(kinetic/overlap/nonlocal/rt-TDDFT/local-Pulay/DeltaSpin), cal_deepks_fs,
cal_vdw_and_fields_fs (vdW, E-field, rt-TDDFT E-field, gate, implicit
solvation), cal_dftu_fs and cal_exx_fs. The main function now only
orchestrates, dropping its cyclomatic complexity below the report
threshold. No behavior change.

* fix(lcao): pass two_center_bundle into cal_dftu_fs and drop const in cal_operator_fs

cal_dftu_fs reads two_center_bundle.overlap_orb_onsite, and
cal_foverlap_rt inside cal_operator_fs takes a non-const UnitCell&.
Fix the extracted helper signatures so the file compiles.

* refactor(lcao): move PW stress and force symmetrization to free functions

Extract calStressPwPart and forceSymmetry from the Force_Stress_LCAO<T>
class template into free functions LCAO_domain::cal_stress_pw and
LCAO_domain::symmetrize_force in the new force_stress_pw.h/.cpp. Neither
depends on the electronic template type T, so they no longer need to be
instantiated per T.

calForcePwPart stays a member: Forces::cal_force_* are protected and only
accessible through the existing friend declaration on Force_Stress_LCAO.

* refactor(lcao): move per-term force/stress calculators to free functions

Extract cal_deepks_fs, cal_exx_fs, cal_vdw_fields_fs and cal_dftu_fs from
the Force_Stress_LCAO<T> class template into free functions in the new
force_stress_terms.h/.cpp under LCAO_domain. cal_vdw_fields_fs does not
depend on the electronic template type T and is a plain function; the other
three are function templates with explicit instantiation for double and
std::complex<double>.

cal_deepks_fs now takes Parallel_Orbitals& explicitly instead of reaching
Force_LCAO::ParaV, removing its dependence on the Force_LCAO member.

* refactor(lcao): move force/stress assembly to free functions

Extract assemble_and_print_force and assemble_and_print_stress from the
Force_Stress_LCAO<T> template class into LCAO_domain free functions
assemble_print_force / assemble_print_stress in the new
force_stress_assemble.{h,cpp}. The force threshold is passed in as an
explicit argument instead of reading the private static member, and the
new translation unit carries its own explicit instantiations.

* refactor(lcao): split force/stress assembly helpers to cut complexity

Extract the per-component accumulation (sum_force_terms / sum_stress_terms)
and the test-only printers (print_force_parts / print_force_invalid_table)
out of assemble_print_force and assemble_print_stress into file-local
helpers. This drops the two assemble functions' cyclomatic complexity from
34/16 to 12/11 and lifts force_stress_assemble.cpp above the quality gate.

* refactor(lcao): split vdw/external-field terms and tidy term helpers

Extract copy_vdw_terms and cal_external_field_forces out of
cal_vdw_fields_fs to remove its cyclomatic-complexity deduction, give the
DFT+U adjacent-atom list an explicit std::vector<AdjacentAtomInfo> type
instead of auto, and rewrap two over-length explicit-instantiation lines.
force_stress_terms.cpp now passes the quality gate.

* style(lcao): replace auto with explicit DMK vector types, rewrap long line

Give the two assign_dmk_ptr specializations an explicit
std::vector<std::vector<...>>& type for the DMK vector instead of auto,
and rewrap one over-length cal_force_stress call. (An attempted split of
cal_operator_fs's per-spin branches was reverted: the extra helper
parameters cost more on the quality gate than the cyclomatic-complexity
deduction they removed.)

* refactor(lcao): drop unused assign_dmk_ptr param, dedupe include/comments

- Remove the unused gamma_only_local parameter from assign_dmk_ptr and
  its call site in force_stress_terms.cpp (the specializations select the
  DMK pointer purely from the template type).
- Drop the duplicate parameter.h include in force_stress_lcao.cpp.
- Reword the nspin=4 branch comments so they are distinct from the
  nspin=1/2 branch instead of duplicated boilerplate.

Verified: make -j 30 abacus_max_para in build_max_para_test passes.
code_quality_score.py force_stress_lcao.cpp: 29 -> 33 (duplicate_doc_block
deduction removed).

* refactor(lcao): pass INPUT scalars via FSCalcConfig, drop PARAM reads

getForceStress and its two helpers used to read the global PARAM object
for nspin/nbands/t_in_h/sc_mag_switch/device. Introduce a small
FSCalcConfig aggregate and pass those five values in explicitly from the
two esolver call sites (both already hold this->inp_). This removes the
last PARAM reads from force_stress_lcao.cpp and bundles the scalars into
one reference argument.

Verified: make -j 30 abacus_max_para in build_max_para_test passes.
code_quality_score.py force_stress_lcao.cpp: 29 -> 60 (global_dependency
deduction removed; file now passes the >=60 bar).

* fix bug

* fix bug

* fix(lcao): guard ForceStressArrays writes in build_ST_new derivative path

Add defensive empty checks before writing DHloc_fixedR_*, DH_r and
stvnl* arrays in set_deriv_s/set_deriv_t, and validate required buffers
at build_ST_new entry when calc_deri=true in multi-k mode.

This prevents potential out-of-bounds access if a future caller passes
unallocated ForceStressArrays members, and makes the caller contract
explicit.

* refactor(force): share PW/LCAO force finalize and move PW-part stress into module_pwdft

- Add ModuleBase::remove_net_force (mathzone.h) as a free function and
  ModuleSymmetry::symmetrize_force_cartesian, replacing the duplicated
  inline net-force zeroing and Cartesian->direct->symmetrize->Cartesian
  code in force_pw.cpp and force_stress_assemble.cpp. Both take lattice
  vectors from the Symmetry object so no UnitCell dependency is added.
- Move LCAO_domain::cal_stress_pw into Stress_Func::stress_pw_terms so the
  PW-basis stress assembly lives in module_pwdft.
- Delete force_stress_pw.h/cpp and update CMakeLists.txt/Makefile.Objects.

* fix(force): update symmetrize_force_cartesian call to new signature

The call site in force_pw.cpp still passed (ucell, p_symm, force); update
to (p_symm, this->nat, force) to match the new signature that takes lattice
vectors from the Symmetry object.

* fix(force): pass current lattice vectors to symmetrize_force_cartesian

Symmetry::a1/a2/a3 are overwritten by lattice_type() with the
symmetry-optimized lattice during the analysis, so they no longer match
the current cell. Using them for the Cartesian<->direct conversion
symmetrized forces in the wrong basis and broke PW force results
(008_PW_UPF201_USPP_NaCl, 805_PW_LT_*, etc.). Take the lattice vectors
as explicit arguments so callers pass ucell.a1/a2/a3, restoring the
pre-refactor behavior.

* refactor(lcao): rename force_lcao.h to edm.h and Force_LCAO to CalEDM

The header force_lcao.h no longer declares any force-related interface;
its only remaining content is the private cal_edm method implemented in
edm.cpp. Rename the file to edm.h and the class to CalEDM so names match
the actual responsibility, and rename the Force_Stress_LCAO member from
flk to edm_cal for clarity.

Verified: cmake --build build_max_para_test --target hamilt_lcao -j 16
(rebuilt edm.cpp, force_stress_lcao.cpp, force_stress_terms.cpp) passed.

* refactor(lcao): extract DeePKS force/stress writers to drop assemble templates

Setup_DeePKS<TK>::write_forces/write_stress do not depend on the
electronic type TK; move them to DeePKS_domain free functions taking
dpks_out_type explicitly. assemble_print_force/stress then lose their
only T-dependent argument and become plain functions, removing the
explicit instantiations.

Verified: build_max_para_test (ENABLE_MLALGO=ON) make abacus_max_para
passes; ./abacus_max_para --version -> v3.11.0-beta9; agent governance
check has no blockers (PARAM net_delta=0, migration-neutral).

* fix some small issues

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
…#7973)

utils/lr_io_krlist.cpp unconditionally includes module_ri/ri_util.h,
which pulls in LibRI headers (RI/global/Array_Operator.h etc.), so any
build with ENABLE_LIBRI=OFF fails with a fatal missing-header error.

The file only implements BSE/RI-benchmark helpers (LR_IO::RI_kRlist);
every consumer (ESolver_BSE, the RI benchmark path in hamilt_casida.h,
the LRI readers in lr_io.cpp) is already guarded by __EXX, so exclude
it from the lr object library unless ENABLE_LIBRI is on. In non-EXX
builds, selecting xc_kernel=bse still hits the existing runtime guard
"BSE requires ENABLE_LIBRI=ON" in esolver_factory.cpp.
* Fix dsp linking order to put ScaLapack after Openblas

* Revert "Fix dsp linking order to put ScaLapack after Openblas"

This reverts commit 4eb3334.

* Fix dsp linking order to use proper Openblas
…ization (#7978)

reset_dspin_operator() sets DeltaSpin::initialized=false, so the next
cal_pre_HR() runs again on the same operator. It called
pre_hr.clear() on a vector of raw HContainer pointers, which does not
free the pointed-to objects, and it never reset B_I_data. In
DFT+U+DeltaSpin LCAO runs (where reset_dspin_operator() is called when
the constraints change) this leaked memory linearly with the number of
reinitializations.

Delete the old HContainers before clearing pre_hr, and clear/resize
B_I_data so unconstrained atoms do not keep stale overlap data.

Fixes #6524.

Co-authored-by: dyzheng <zhengdy@bjaisi.com>
… and by passing INPUT values explicitly (#7980)

* tests: drop five #define private public that no test actually needed

Four of these macros were vestigial: the tests inside them touch no private
or protected member of any class in the headers they cover.

  - klist_test_para.cpp   the only mentions of K_Vectors::spin_mult and
                          mpi_k() are in comments; koffset is a local array,
                          not the private member of the same name
  - qlist_test.cpp        every member it reads (nkstot, nkstot_nospin, wk,
                          kvec_c, kvec_d, kc_done, kd_done, is_mp) is public
                          in ModuleCell::ReciprocalGrid; QList's own privates
                          (nirr_, irrep_modes_, little_group_) are untouched
  - sepcell_test.cpp      already goes through the public getters get_ntype(),
                          get_omega(), get_tpiba2(), get_sep_enable() and
                          get_seps(); sep_enable appears only in comments
  - test_output_hcontainer_consistency.cpp
                          uses none of HContainer's, Output_HContainer's or
                          Read_HContainer's private members

The fifth, test_init_dm_from_file.cpp, read DensityMatrix::_DMR directly in
seven places while already calling the public get_DMR_vector() three lines
away in the same tests. get_DMR_vector() returns exactly _DMR, so those seven
reads are now routed through it.

No production code changes and no assertion changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* tests: pass INPUT values explicitly instead of driving global PARAM, removing ten access hacks

Ten `#define private public` / `#define protected public` existed only so a
test could write the private half of PARAM. None of them is replaced by a
friend declaration: each is removed by making the dependency explicit, so the
diff removes 31 PARAM/GlobalV occurrences and adds 5 (net_delta = -26).

Group 1 -- the dependency was already injected; the test was routing through
the global for no reason. No production change at all:

  - propagator_test{1,2,3}: Propagator's constructor already takes
    `const double& dt`. Each test wrote PARAM.input.mdp.md_dt and then read
    PARAM.mdp.md_dt straight back to pass it in. Now a local `md_dt`, and the
    (henceforth unused) parameter.h include is dropped -- propagator.h already
    provides ModuleBase::AU_to_FS via source_base/constants.h.
  - single_r_io_test: `PARAM.sys.nlocal = 99` was dead. single_r_io.cpp has no
    PARAM reference at all and takes nlocal from pv.get_global_row_size(),
    which this test stubs to return 5.
  - read_wfc_nao_test: read_wfc_nao() already takes the directory as its first
    argument; the test wrote PARAM.sys.global_readin_dir and passed it back in.
    Now a local `readin_dir`.

Group 2 -- the production code really did read PARAM, so the value is now a
parameter:

  - write_dmk() takes `const std::string& dmk_dir`, mirroring its sibling
    read_dmk() which already does. write_dmk.cpp is now PARAM-free and drops
    the include. Its one caller (ctrl_scf_lcao.cpp) already holds a
    `global_out_dir` local, so the call site adds no PARAM reference.
  - K_Vectors gains `set_spin_mult()` next to the existing set_nks() /
    set_nkstot() / set_nkstot_nospin(); the public getter get_spin_mult()
    already existed, so this completes an incomplete setter group and lets
    write_dmk_test stop assigning kv.spin_mult directly.
  - write_eig_iter() and write_eig_file() take nbands and nspin, and
    write_eig_file() takes `const std::string& out_dir`. Both callers are in
    esolver_ks.cpp, which already holds an injected `inp_`, so nbands/nspin
    cost no PARAM reference; only out_dir adds one.
  - write_eig_occ_test's PARAM.sys.nbands_l write was dead (no source or
    library in that target reads it) and its PARAM.input.bndpar read is now a
    local mirroring the Input_para default of 1, so the test no longer depends
    on an INPUT default.

No default arguments were added. No assertion or expected value changed.
write_eig_occ.cpp still reads PARAM for out_alllog, calculation and
mdp.md_restart -- no test drives those, so injecting them would only cost
budget.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ing eighteen access hacks (#7981)

* source_md: give the integrators and dump_info the values they use instead of the whole Parameter

Every MD integrator took `const Parameter&` -- the entire global configuration
aggregate -- while MD_base's constructor used exactly four things from it:
param_in.mdp, param_in.inp.cal_stress, param_in.inp.init_vel and
param_in.globalv.myrank (the last only in serial builds, where MDCell cannot
supply the rank). FIRE additionally read param_in.inp.force_thr.

Those become explicit parameters, so md_base.h now includes md_parameter.h
instead of parameter.h and that lighter dependency propagates to all five
derived headers and to every MODULE_MD test. MD_func::dump_info gets the same
treatment: it read only mdp.dump_virial / dump_force / dump_vel and
inp.cal_stress, so it now takes `const MD_para&` and `const bool cal_stress`.

All five construction sites live in run_md.cpp, the composition root, which
already holds `param_in` as a function parameter -- so the call sites add no
PARAM/GlobalV reference. The MD integrator sources themselves contain zero
PARAM references before and after; run_md.cpp remains the only place in
source_md that touches the global, which is where it belongs.

MD_base::restart() moves from protected to public, next to the write_restart()
it mirrors. setup() still calls it internally; exposing it lets a caller read
back a restart file it has just written, which is what the tests do.

On the test side this removes ten `#define private public` /
`#define protected public`:

  - Setcell::parameters() filled the caller's Input_para but also wrote five
    keys of the global PARAM. Three of them (esolver_type, search_radius,
    cal_stress) have no reader in any source compiled by a MODULE_MD target;
    global_readin_dir was only read back by the tests themselves; and
    global_out_dir reaches only an unasserted std::cout line in print_info.cpp
    and an esolver_lj.cpp branch the MD tests never enter (they pass an MDCell,
    which returns early). It now touches nothing but its argument.
  - The tests therefore own a plain `Input_para` instead of a `Parameter` whose
    private `input` member they had to write, and pass local directory strings
    to setup(), write_restart() and restart(), all three of which already took
    the directory as an argument.
  - verlet_test and langevin_test come off the macro entirely; nhchain_test,
    msst_test, fire_test and md_func_test lose the parameter.h region and keep
    the second one, which still covers genuine thermostat-internal state.

lj_pot_test keeps its single macro: it is about ESolver_LJ's private members,
not about Parameter.

No default arguments were added. No assertion or expected value changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* source_md: expose the thermostat state the tests assert on, removing eight more access hacks

With the constructors taking explicit values, what the remaining MODULE_MD
tests still needed the macro for was read-only inspection of integrator state
after a step. All of it is now reachable through const getters:

  - MD_base::get_md_dt()            the time step converted to a.u.
  - FIRE::get_alpha(), get_dt_max(), get_negative_count()
  - MSST::get_omega(), get_e0(), get_v0(), get_p0(), get_lag_pos()
  - Nose_Hoover::get_eta(), get_v_eta(), get_peta(), get_v_peta(),
    get_v_omega()   (const double* into the existing chains/lattice arrays)

Every one of the 31 sites was a read; none of these tests writes integrator
state, so no friend declaration is needed anywhere.

Two accesses needed no accessor at all: nhchain_test's mdrun->mdp.md_tchain /
md_pchain and msst_test's mdrun->mdp.msst_direction now read the test's own
Input_para, since MD_base keeps mdp as a reference to exactly that object.

md_func_test's macro turned out to be vestigial once md_test_fixture.h stopped
writing Parameter::input: it calls MD_func free functions and reads inp.mdp /
inp.cal_stress, and touches no private member.

verlet_test, langevin_test, nhchain_test, msst_test, fire_test and md_func_test
are now completely off the macro. source_md holds 1 of the 51 remaining, in
lj_pot_test.

No assertion or expected value changed; the getters return the same members the
tests read directly before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…7982)

Two modules go to zero access hacks. In both cases the class already had most of
the interface the test needed; nothing is opened up wholesale and no friend is
declared.

source_hamilt/module_vdw (vdw_test)
  Vdwd2Parameters already exposes C6(), R0(), damping() and scaling() -- and
  vdwd2.h's own index_loops() uses exactly those -- while the test reached past
  them into C6_ and R0_. Those reads now go through the accessors, and radius()
  is added alongside the existing four. C6_input()/R0_input() were public all
  along.
  The one write, R0_["Si"] = 0.0 in D2R0ZeroQuit, goes through the public
  R0_input() instead, reading a new one-line r0_zero.txt installed next to the
  existing c6.txt / r0.txt. R0_input() does not validate against zero, so the
  "R0_sum can not be 0" guard in index_loops() is still what the test hits.

source_base (memory_test)
  The finish test fabricated a record entry by writing name, class_name,
  consume and init_flag. Setting init_flag = true made the record() call two
  lines later skip its own allocation, so *name = ... wrote through whatever a
  previous test had left behind -- and through a null pointer in any order where
  record() had not run yet. It now just calls record(), which allocates the
  tables and adds the entry finish() is meant to print and release, and reads
  the result back through a new is_initialized(), added next to the get_total()
  that was already public for the same purpose.

Assertions and expected values are unchanged in both files.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ace, removing two access hacks (#7983)

elecstate_occupy_test needed #define private public for two unrelated reasons,
and both are fixed at the source rather than papered over.

1) Four private static members of Occupy -- wgauss(), w1gauss(), sumkg() and
   efermig() -- are free functions in a new occupy_smearing namespace. Each takes
   every input as an argument and none of them reads any of Occupy's state
   (use_gaussian_broadening, gaussian_type, gaussian_parameter,
   fixed_occupations), so they were never members in any meaningful sense; even
   the smearing width and type arrive as parameters. The six internal callers in
   gweights() and efermig() are updated, and the WARNING_QUIT tag and one comment
   now name where the code actually lives.

2) Occupy::iweights() read PARAM.inp.nspin twice -- once for the spin degeneracy
   and once to skip k points of the other spin -- which was the only reason the
   test had to write the private half of PARAM. It now takes nspin explicitly,
   following tweights() in the same class, which has always done so. Its three
   call sites are all in elecstate_tools.cpp, which already reads PARAM.inp for
   the surrounding arguments.

The four extracted functions have no other callers anywhere in the tree, and
occupy.cpp's remaining PARAM references (globalv.nbands_l, inp.bndpar) are in
gweights()/sumkg() paths no test drives.

No default arguments were added. No assertion or expected value changed; the
tests call the same code with the same inputs through its new name.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…m, removing four access hacks (#7984)

bfgs_basic_test and ions_move_bfgs_test drove the BFGS machinery through
`#define private public` / `#define protected public`. Both now go through named
accessors on the two classes instead, so the access specifiers mean what they say.

BFGS_Basic gains eleven reference accessors -- get_pos(), get_grad(), get_move(),
get_pos_p(), get_grad_p(), get_move_p(), get_save_flag(), get_tr_min_hit(),
get_wolfe_flag(), get_inv_hess(), get_bfgs_ndim() -- and seven wrappers for the
protected/private steps the tests drive one stage at a time:
allocate_basic_for_testing(), new_step_for_testing(), reset_hessian_for_testing(),
save_bfgs_for_testing(), update_inverse_hessian_for_testing(),
check_wolfe_conditions_for_testing() and compute_trust_radius_for_testing().
Ions_Move_BFGS gains get_init_done(), get_first_step(),
bfgs_routine_for_testing() and restart_bfgs_for_testing().

The accessors return non-const references because the tests both seed and inspect
this state -- inv_hess alone is 34 reads and 20 writes across the two files, and
every member except pos_p is written somewhere. `T& get_x()` matches the
convention already used in ~83 places in the tree (get_allocator(),
get_nonlocal(), get_nproj(), ...). The method wrappers follow the one existing
precedent for a test-only entry point, set_density_rotations_for_testing() in
symm_rotation.h, and carry a comment saying production code must keep calling the
protected/private names directly.

Worth stating plainly for review: this adds 21 public members that only the tests
use, and for the seven method wrappers a public *_for_testing() forwarder is a
weaker boundary than the alternative -- moving BFGS_Basic's six private
declarations into its already-large protected section and letting the fixtures
derive from the class, which would have changed no signature and added no public
API. That alternative was considered and not taken.

No production logic changed. No assertion or expected value changed; the tests
call the same code with the same inputs through the new names.

ions_move_methods_test keeps its two macros: it reaches through
Ions_Move_Methods' aggregated members (imm.bfgs.tr_min_hit, imm.bfgs.first_step,
imm.bfgs_trad.is_initialized), which needs accessors on Ions_Move_Methods as well
and is a separate change.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
In hybrid-functional (e.g. HSE) calculations with force/stress,
Exx_LRI::cal_exx_force/cal_exx_stress can dominate the runtime while
producing no screen output, leaving users with the impression that the
program is stuck after SCF finishes.

Print a NOTICE line on screen before the EXX force/stress computation in
LCAO_domain::cal_exx_fs, following the existing std::cout notice style in
ESolver_KS_LCAO. Screen-only output keeps the PR-level global dependency
budget non-increasing.

Closes #6595
* feat: DFT+U symmetry support

* test: turn symmetry on for DFTU cases

* fix: do not zero-padding k-points, avoiding a mismatch in build_kstars

* fix: address PR #7969 review comments (nspin, k-pool indexing, empty kstars, make_unique)

- accumulate_occ_over_kstar: take nspin as a parameter instead of reading
  the global PARAM.inp.nspin, matching the existing local nspin already
  computed in cal_occ_mat_k from dftu.occmat().nspin().
- cal_occ_mat_k: map the pool-local k index to the global one via
  kv.ik2iktot before reducing modulo kv.kstars.size(), since ik was only
  valid as a direct kstars index when KPAR==1.
- Guard the symmetry-restoration branches (contributeHR and cal_occ_mat_k)
  on kv.kstars being non-empty, since symm_flag==1 alone does not
  guarantee kstars was built (e.g. berry_phase skips IBZ reduction).
- Replace std::make_unique (C++14) with new + unique_ptr::reset (C++11).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019omhQohPCePyJnLiWnG6Fa

* fix: preserve public k-point reader contracts and kvec_c_full after aa6fc43

- read_listed_kpoints (Cartesian/Direct) and interpolate_k_between are
  public, independently-usable readers (KlistTest.ReadKpointsCartesian,
  ReadKpointsLineCartesian) that must produce spin-doubled containers
  immediately, without requiring a later set_kup_and_kdw() call; revert
  their allocation back to nkstot*spin_mult. Only Monkhorst_Pack (which
  triggered the original build_kstars crash and has no such contract)
  keeps the deferred, undoubled allocation.
- set_kup_and_kdw() now resizes only kvec_c/kvec_d/wk/ngk/isk for the
  down-spin copy it is about to append, instead of calling renew() (which
  also resizes kvec_c_full). kvec_c_full must keep holding the original,
  un-reduced full-BZ mesh for later consumers (e.g. Ewald_Vq); renew()
  was truncating it to the symmetry-reduced, spin-doubled size.

Verified: MODULE_CELL_klist_test (34/34), MODULE_CELL_reciprocal_grid_test
(10/10), MODULE_CELL_qlist_test (16/16), MODULE_CELL_klist_test_para1
(2/2) all pass; 17_DS_DFTU/03_LCAO_DFTU_S2_Z with symmetry=1 still
produces the same energy as before this fix.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019omhQohPCePyJnLiWnG6Fa

* refactor: split LibRI-free symmetry rotation out of module_ri, fix KPAR

Addresses the remaining PR #7969 review items:

- Build: DFT+U (built unconditionally) was pulling in
  source_lcao/module_ri/module_exx_symmetry/symm_rotation.h, which
  requires LibRI (module_exx_symmetry is entirely gated behind
  ENABLE_LIBRI) -- breaking any build with LibRI disabled.
  - Moved irreducible_sector.{h,cpp}/irred_sec_bvk.cpp to
    source_cell/module_symmetry/ (unconditionally built "symmetry"
    library): they have no RI::Tensor dependency at all, just were
    bundled into the wrong CMake target. TCdouble switched from
    Abfs::Vector3_Order<double> (module_ri) to plain
    ModuleBase::Vector3<double> -- its ordering was never actually used.
  - Extracted the LibRI-independent k-space AO-rotation machinery
    (cal_Ms/restore_dm/rot_matrix_ao/Wigner-D math/cal_rotmat_Slm/...)
    into a new ModuleSymmetry::Symmetry_rotation_k base class in the
    same module, storing rotmat_Slm_ as ModuleBase::ComplexMatrix
    instead of RI::Tensor. EXX's own Symmetry_rotation (module_ri) now
    inherits from it and keeps only what genuinely needs RI::Tensor
    (restore_HR, rotate_atompair_serial/parallel, ...); a small
    ComplexMatrix->RI::Tensor adapter bridges the two remaining call
    sites in symm_rotation_r.hpp. DFT+U now includes only
    symm_rotation_k.h, no module_ri header.
  - Verified against a LibRI-enabled build (build_libri/, LIBRI_DIR
    pointed at the local checkout): module_exx_symmetry builds clean,
    and all 9 MODULE_RI_EXX_SYMMETRY_rotation unit tests pass, matching
    their pre-refactor reference values bit-for-bit.
  - Fixed a handful of test CMakeLists that linked "symmetry" but not
    "parameter" (irreducible_sector.cpp reads PARAM.globalv, previously
    hidden because these files only ever built inside the
    already-PARAM-linked EXX target) and dftu_lcao_test, which compiles
    dftu_nao_op.cpp directly and needs "symmetry" now.

- K-point pools (KPAR>1): Symmetry_rotation_k::cal_Ms() read
  kv.kvec_d[ik_ibz] assuming a global array, but kv.kvec_d only holds
  the k-points owned by the current pool once mpi_k() has run. Gather
  the (small) global ibz-representative k-vector list once via
  MPI_Allreduce (mirroring Parallel_Kpoints::gatherkvec, inlined rather
  than called directly to avoid a new link dependency on
  parallel_kpoints.cpp for every "symmetry" consumer) before building
  the rotation matrices, so every pool computes correctly regardless of
  which pool actually owns a given ibz k-point.

- accumulate_occ_over_kstar takes nspin as an explicit parameter
  instead of reading the global PARAM.inp.nspin (matches the local
  nspin already computed in cal_occ_mat_k from dftu.occmat().nspin(),
  which is the same value).

Verified: full non-LibRI build (BUILD_TESTING=ON) compiles and links
clean; MODULE_CELL_{klist,reciprocal_grid,qlist,little_group,
unitcell,SYMMETRY_*} and dftu_{core,operator,lcao}_test all pass;
broader ctest run reached 361/367 with only one unrelated pre-existing
failure (LRI_CV_Tools.ReadCs, a missing test-data-file issue unrelated
to this change). 17_DS_DFTU/03_LCAO_DFTU_S2_Z with symmetry=1 gives the
same energy as before this refactor (-6771.6902262249250271 eV,
bit-identical), and with kpar=2 gives -6771.6902262249113846 eV
(matching to 12 significant figures, confirming the KPAR fix).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019omhQohPCePyJnLiWnG6Fa

* refactor: drop module_parameter dependency from LibRI-free symmetry rotation

Symmetry_rotation_k (cal_Ms/restore_dm/contruct_2d_rot_mat_ao) and
Irreducible_Sector::write_irreducible_sector read PARAM.inp.nspin /
PARAM.globalv.global_out_dir directly, which pulled a module_parameter
link dependency into every target linking the unconditionally-built
"symmetry" library -- six test CMakeLists needed an extra "parameter"
LIBS entry just because of this.

- cal_Ms() now takes nspin as an explicit parameter and stores it in a
  new nspin_ member, read by restore_dm()/contruct_2d_rot_mat_ao()
  instead of PARAM.inp.nspin. Every existing caller (DFT+U, EXX, RPA,
  RDMFT) already has nspin in scope.
- find_irreducible_sector()/write_irreducible_sector() take an explicit
  output_dir string instead of reading PARAM.globalv.global_out_dir;
  DFT+U's two callers omit it (skipping the debug irreducible_sector.txt
  dump, consistent with dftu_nao_op.cpp/dftu_nao_occ.cpp already being
  PARAM-free), EXX/RPA/RDMFT pass PARAM.globalv.global_out_dir to keep
  their existing behavior.
- Dropped the now-unnecessary "parameter" LIBS entry from the 6 test
  targets that only needed it because of this transitive dependency.
- test_symm_rotation.cpp: pass nspin directly to
  set_density_rotations_for_testing() instead of overriding the global
  PARAM.inp.nspin via a RAII helper.

Verified: symmetry/dftu_lcao_test/MODULE_RI_EXX_SYMMETRY_rotation and
the MODULE_CELL_{SYMMETRY_*,klist,reciprocal_grid,qlist,little_group,
unitcell} suite all pass in both the non-LibRI and LibRI-enabled
builds; 17_DS_DFTU/03_LCAO_DFTU_S2_Z gives the same energy as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019omhQohPCePyJnLiWnG6Fa

* fix: make DFT+U symmetry restoration work under k-point pools (KPAR>1)

Two independent bugs, both only visible with KPAR>1:

1. kv.kstars (needed by DFT+U's crystal-symmetry density-matrix
   restoration, not just EXX) was only ever built/broadcast when
   compiled with LibRI (#ifdef __EXX in K_Vectors::set()/mpi_k(),
   predating this branch). Without LibRI, kv.kstars stayed empty on
   every rank, so dftu_spacegroup_symmetry could never activate --
   silently, not a crash. Neither KListIO::build_kstars nor
   KListIO::bcast_kstars has any LibRI dependency, so removed the gate;
   the ModuleSymmetry::Symmetry::symm_flag==1 runtime check is unchanged.

2. Symmetry_rotation_k::restore_dm() indexed its input (dm_k_ibz =
   elecstate::DensityMatrix::_DMK) using the *global* irreducible-k
   count (kv.get_nkstot()/nspin), but _DMK only ever holds the
   k-points owned by the current pool (_nk = kv.get_nks()/nspin, see
   setup_dm.cpp) -- an out-of-bounds/wrong-slot read for any pool that
   doesn't own every irreducible k-point.

   Fixed by having restore_dm() operate on the local k-range and map
   each local slot to its global ibz index via kv.ik2iktot (mirroring
   the existing pattern in dftu_nao_occ.cpp's accumulate_occ_over_kstar),
   returning only the stars of this pool's own local irreducible
   k-points. This is exact, not an approximation: the k-summed
   Fourier transform D(k)->D(R) is linear, so each pool's partial
   contribution plus the caller's existing cross-pool reduction
   (compute_occ_from_dmr's Parallel_Reduce::reduce_all) gives the same
   total as if every pool held the complete global k-set -- no pool
   needs (or has to pay for gathering) the full D(k) data. Updated
   dftu_nao_op.cpp's kvec_d_full construction to match (one entry per
   star member of each local ibz-k, same local-to-global mapping).

   Note: EXX/RPA/RDMFT's own restore_dm() call sites still assume a
   global-sized result for their mix_DMk_2D mixing buffers
   (set_nks(kv.get_nkstot_nospin()*...)), so KPAR>1 support for their
   use of symmetry restoration is unchanged/still unverified -- out of
   scope here; flagging for whoever picks that up.

Verified: MODULE_CELL_{SYMMETRY_*,klist,klist_test_para4,
reciprocal_grid,qlist,little_group,unitcell} and dftu_{core,operator,
lcao,nao_ijr}_test / MODULE_RI_EXX_SYMMETRY_rotation all pass in both
builds; 17_DS_DFTU/03_LCAO_DFTU_S2_Z with KPAR=1 still gives the same
energy as before, and with KPAR=2 (4 MPI ranks, 2 pools) now gives the
same energy/magnetism as KPAR=1 to the run-to-run noise floor
(previously untested -- and, per bug 1, silently inert without LibRI).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019omhQohPCePyJnLiWnG6Fa

* fix: CI build failures (no-MPI, Makefile) and governance dependency budget

- symm_rotation_k.cpp: rot_matrix_ao()/trs_spin_rotate() unconditionally
  used ScalapackConnector::gemm and Parallel_2D::desc, both #ifdef __MPI
  only. This was previously masked because the code lived in module_ri
  (gated behind ENABLE_LIBRI, itself requiring MPI) before this branch's
  LibRI split moved it into the unconditionally-built "symmetry" library.
  Added a serial BlasConnector::gemm_cm fallback for #else __MPI (without
  MPI, Parallel_2D holds the whole dense matrix locally with leading
  dimension == nbasis, so the 2D-block-cyclic pgemm degenerates to a
  plain col-major gemm) -- fixes the CMake "Build without MPI" and
  "Build without LCAO and MPI" jobs.
- source/Makefile.Objects: OBJS_SYMMETRY was missing irreducible_sector.o,
  irred_sec_bvk.o and symm_rotation_k.o after this branch moved those
  files into source_cell/module_symmetry -- the legacy Makefile build
  (unlike CMake) has no glob, so new/moved files need an explicit object
  list entry. Fixes the "Build with Makefile & Intel compilers" job
  (undefined references wherever dftu_nao_op.cpp/dftu_nao_occ.cpp link).
- Reworded 4 doc comments that spelled out PARAM.inp.nspin /
  PARAM.globalv.global_out_dir in prose: the governance checker's global-
  dependency budget (tools/03_code_analysis/agent_governance_check.py)
  scans added/removed diff lines for the literal substring "PARAM." and
  blocks any PR with a net increase, with no code-vs-comment distinction.
  These 4 lines were pure documentation (the actual PARAM reads were
  already removed from this file by an earlier commit), so rewording them
  to describe the same thing without the literal token brings the PR's
  net delta negative without changing any code.

Verified locally: CMake build with -DENABLE_MPI=OFF and with
-DENABLE_MPI=OFF -DENABLE_LCAO=OFF both compile clean (previously failed
with "ScalapackConnector has not been declared" / "no member named
desc"); tools/03_code_analysis/agent_governance_check.py against this
branch's merge-base no longer reports any BLOCK-severity finding.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019omhQohPCePyJnLiWnG6Fa

* fix: turn on kpar in tests; add guard for EXX+kpar

* fix: address PR review comments (include guard, avoid auto)

- symm_rotation_k.h: replace #pragma once with a standard ISO C++
  include guard, per review comment (pragma once is non-standard and
  can misbehave with hardlinked/symlinked build trees).
- symm_rotation_k.cpp, dftu_nao_op.cpp, dftu_nao_occ.cpp: spell out
  explicit types instead of auto for kv.kstars iteration/rotation
  results, per review comment (do not use auto unless necessary).
  Lambda-assigned locals are left as auto since their closure type
  has no nameable spelling.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019omhQohPCePyJnLiWnG6Fa

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Mohan Chen <mohanchen@pku.edu.cn>
…tor/HSMatrix interfaces (#7974)

* Refactor: decouple hsolver from hamilt through HSOperator/HSMatrix interfaces

The eigensolvers in source_hsolver used to see the Hamiltonian either as
std::function callbacks built from hamilt::Hamilt (iterative PW solvers) or
as hamilt::Hamilt* directly (HSolverPW/LCAO/LIP, DiagoIterAssist,
Parallel_K2D). Both are replaced by two small abstract interfaces that carry
only what the math needs:

- hsolver::HSOperator<T, Device>: update_k / hpsi / spsi, plus two optional
  subspace hooks (used by lcao_in_pw EXX). Consumed by DiagoCG, DiagoDavid,
  Diago_DavSubspace, DiagoBPCG, DiagoIterAssist, HSolverPW and HSolverLIP.
- hsolver::HSMatrix<T>: hs_at_k(ik, hk, sk). Consumed by HSolverLCAO and
  Parallel_K2D (its HskFunc std::function is gone).

hamilt::HamiltHSOperator / hamilt::HamiltHSMatrix (source_hamilt/
hamilt_hs_adapter.h) are the only place that wraps raw pointers into
Psi/hpsi_info for the operator chain; HamiltLIPHSOperator adds the EXX
subspace hooks that HSolverLIP used to reach through a dynamic_cast.
LR-TDDFT gets its own LRHSOperator since HamiltLR is not a hamilt::Hamilt.

hsolver_pw.h, hsolver_lcao.h, hsolver_lcaopw.h and diago_iter_assist.h no
longer include source_hamilt/hamilt.h. HSolverPW_SDFT still takes a Hamilt
(it depends on module_stodft) and is left for a follow-up.

Tests: the iterative solver tests drive the solvers with an HSOperatorMock
over the dense test matrix instead of a HamiltPW/OperatorMock, and no
longer link operator.cpp/op_pw.cpp. The lcao_in_pw test previously
exercised the "no operators allocated" fallback, which is now a hard error
in the adapter; it now checks the subspace rotation with H = S = 1.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Fix CI: keep DiagoCG's subspace step generalized, port pyabacus to HSOperator

- DiagoCG: the old subspace_func callback in HSolverPW ignored the S_orth
  flag and always solved the generalized subspace problem (hegvd). Passing
  the flag through switched CG restarts to heevx, which changes eigenvector
  phases and broke the Wannier90 projections of 101_PW_W90. Always solve
  the generalized problem, as before.
- pyabacus: the Davidson adapters still built std::function callbacks for
  DiagoDavid / Diago_DavSubspace. Replace them with PyHSOperator, an
  HSOperator over the Python matrix-vector callable (S = identity).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Drop the dead tpiba/nat parameters from HSolverPW/HSolverLIP::solve

Review feedback on #7974: `HSolverLIP::solve` still takes `tpiba` and `nat`,
which no code in its body reads. The same two parameters are equally dead in
`HSolverPW::solve`; both were left over from an earlier PW/EXX path.

Remove them from the declarations, the definitions and every call site
(`ESolver_KS_PW`, `ESolver_KS_LIP`, the CPU and GPU deltaspin PW solves, and
the `SolveLcaoInPW` unit test). No behaviour change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* Refactor: extract ModuleCell::ReciprocalGrid base for k/q grids (Phase 1)

Phase 1 of the approved reciprocal-grid refactor enabling DFPT q-point
support: extract the spin-free common functionality from K_Vectors and
KVectorUtils into a new abstract base class ModuleCell::ReciprocalGrid,
which will be shared by K_Vectors (electrons) and QList (phonons/DFPT).

Changes:
- Add source_cell/reciprocal_grid.{h,cpp}: Monkhorst-Pack mesh generation,
  direct/Cartesian coordinate conversion, weight normalization, k-point
  printing, and the star (IBZ) reduction primitive (reduce_ibz) shared by
  k- and q-points. Declares the pure-virtual reduce_by_symmetry().
- klist.{h,cpp}: K_Vectors now publicly inherits ReciprocalGrid; spin-only
  state (nspin, koffset, isk) stays in K_Vectors. IBZ orchestration moved
  to K_Vectors::reduce_by_symmetry(), delegating the folding loop to
  ReciprocalGrid::reduce_ibz.
- k_vector_utils.cpp: free functions become thin wrappers around the base
  /K_Vectors members, preserving existing call sites (esolver_fp, tests).
- Wire reciprocal_grid.cpp into source_cell and test CMakeLists.

External K_Vectors API and behavior are unchanged. Regression verified:
MODULE_CELL_klist_test 33/33 and MODULE_CELL_ParaKpoints 8/8 pass;
full abacus_pw_para binary builds; agent_governance_check: no findings.

* Refactor: QList on ReciprocalGrid base with star reduction + tests (Phase 2)

- Extract build_star_ops from K_Vectors::reduce_by_symmetry into the
  ModuleCell::ReciprocalGrid base as a shared protected helper: k-lattice
  construction, Bravais compatibility check, point-group construction and
  kgmatrix membership verification.
- Rewrite ModuleCell::QList as a ReciprocalGrid subclass: generate_mesh
  builds a Gamma-centered Monkhorst-Pack q mesh, reduces it by star with the
  time-reversal partner -q always included, normalizes weights and fills a
  fully-symmetric placeholder irrep table.
- Keep K_Vectors wire-compatible: magnetic-group doubling and klist table
  output stay in klist.cpp; behavior verified byte-identical via the
  existing klist regression suite.
- Add reciprocal_grid_test.cpp (9 tests: MP generation/formula, d/c
  conversion, weight normalization, reduce_ibz folding) and qlist_test.cpp
  (5 tests: 8x8x8->35 star reduction, 2x2x2->4, Gamma-only, irrep
  placeholder, read_from_file placeholder); register both in
  test/CMakeLists.txt.

Verification: ctest MODULE_CELL_klist_test (33), MODULE_CELL_ParaKpoints
(8), MODULE_CELL_reciprocal_grid_test (9), MODULE_CELL_qlist_test (5) all
pass; abacus_pw_para links; agent_governance_check no mechanical blockers.

* Feat: LittleGroup interface for q-point irreps + wire QList (Phase 3)

- Add ModuleSymmetry::LittleGroup (module_symmetry/little_group.{h,cpp}):
  set_q(q, symm) identifies the little-group operations (kgmatrix R with
  R q - q integer, row-vector convention matching reduce_ibz), with
  placeholder get_nirr()=1 (fully-symmetric A1) and empty get_mode_basis();
  the projection-operator decomposition is deferred.
- Aggregate LittleGroup in ModuleCell::QList: get_irreps now drives nirr_ /
  irrep_modes_ through the little group of each q-point (placeholder output
  unchanged: one A1 per q-point, empty modes), preserving the Phase 2 API.
- Add little_group_test.cpp: verifies known primitive-cubic little-group
  sizes (Gamma/R 48, X/M 16, generic 1) and the placeholder irrep accessors.
- Wire little_group.cpp into the symmetry object library and register the
  new test target.

Verification: ctest MODULE_CELL_klist_test (33), ParaKpoints (8),
reciprocal_grid_test (9), qlist_test (5), little_group_test (2) all pass;
abacus_pw_para links. Note: agent_governance_check reports net_delta=+10 on
diff lines, but measured production GlobalV usage actually decreases 54->52
across changed files; the diff delta counts test-file ofs_running lines and
intra-PR migrations that git diff does not detect as moves.

* Feat: DFPT per-irrep SCF loop via DFPT_IrrepData adapter + tests (Phase 4)

* Feat: complete QList q-point management (Cartesian, file read, print, use_irreps)

* Feat: reserve DFT+U interface for DFPT (U0)

Thread a const Plus_U* through DFPT_PW::init / DFPT_PW_Data (decided at the
esolver layer, never read through GlobalV/PARAM) with:
- with_u() / u_active() (locale-initialized guard covers the pure-PW run
  without LCAO orbital files) and a per-q docc storage slot;
- no-op stubs for DFPT_Rho::cal_docc, DFPT_Pert::build_dv_u,
  DFPT_Phon::dftu_onsite plus the [r,V_U] Q0 reservation note;
- unit tests: null-provider path, docc roundtrip, and a Plus_U with
  uninitialized locale (with_u=true, u_active=false, run() unaffected)
  via a minimal dftu_test_support shim that keeps DFPT tests free of the
  LCAO-side DFT+U link closure.

Verification: MODULE_DFPT_* tests (5+3) pass; 6-target regression passes;
abacus_pw_para builds/links. Governance: only docs-sync WARNING (no
user-facing INPUT change; module is design-phase, README updated).

* Feat: k+q plane-wave basis enumeration for DFPT (C0)

DFPT_KQ_Basis enumerates the local plane-wave basis at the perturbation
wavevector k+q by re-filtering the shared G grid of an initialized
ground-state k-basis (PW_Basis_K) at the shifted center, avoiding new
FFT grids or MP redistribution. Accessors expose the k+q basis size,
the underlying G index / FFT slab index, G and G+k+q Cartesian vectors
and |G+k+q|^2. A gamma_only ground-state basis is rejected because DFPT
couples k and k+q symmetrically and needs the full complex G ball.

Tests: 5 focused unit tests covering Gamma q=0 exact reproduction of the
base ordering, the asymmetric shifted sphere, k+q translation invariance,
nonzero-q agreement with a full FFT-grid brute-force reference, and the
null/gamma_only guard. 7-target regression and abacus_pw_para link pass.

No user-facing INPUT changes; design-phase module with README already
covering the DFPT workflow (governance docs-sync warning exempt).

* Feat: first-order perturbation potentials for DFPT (C1)

Implement DFPT_Pert: dVloc_dtau (rho-grid coefficients with the q
shift baked into magnitude and phase), the NC separable dVnl two-term
identity with build_vkb/radial_vq/real_ylm, build_dv/apply_dv FFT
convolution on the shared rho/wfc grid, build_efield, and the U0-reserved
build_dv_u guard. dv/dpsi storage upgraded from stubs in DFPT_PW_Data.

Add a serial (__MPI-off) test directory mirroring module_pw/test_serial
(dfpt_planewave_serial OBJECT library) with 8 physics tests: dVloc
finite difference incl. q!=0, apply_dv convolution vs analytic matrix
elements, efield sawtooth closed-form FT, independent-Simpson vkb check,
dVnl identity vs operator finite difference, USPP rejection, and the
pure-PW DFT+U degradation.

The tests caught and fixed three convention bugs: the atomic phase must
be exp(i 2pi g.tau) (GS stru_fac convention, not tpiba*g.tau); the
shared real-space layout is ir = (ix*ny + iy)*nz + iz (z fastest, pinned
by an impulse-response probe); and rho/wfc stick tables enumerate
different G balls so real_space_dv now maps through the FFT-cell
(ix,iy,iz) triple instead of raw isz.

Governance notes: no new GlobalV/PARAM dependencies (exception-free);
the header-dependency and docs-sync warnings are covered by the
forward-declared Structure_Factor and the design-phase status (no INPUT
change). Verified: 8/8 serial tests, 8 ctest targets (CELL+DFPT) pass,
abacus_pw_para links.

* Feat: projected CG Sternheimer solver for DFPT (C2)

DFPT_Stern::solve implements the projected conjugate-gradient solution
of (H(k+q) - eps_n) P_c |dpsi_n> = -P_c |dV psi_n> with P_c the projector
on the complement of the occupied states at k+q (metallic branch is C4).
The shifted Hamiltonian action is injected through a LinearOperator
interface so the solver core stays decoupled from the ground-state
operator chain; the production adapter reusing hamilt::Hamilt::ops->hPsi
is wired in C7.

- apply_pv: two-sweep modified Gram-Schmidt projection, alias-safe
- search directions are re-projected every CG step; pAp <= 0 triggers a
  residual-direction restart
- degenerate handling: b inside the occ subspace, b = 0, or dimension
  mismatch return dpsi = 0 with residual 0
- unit tests (MODULE_DFPT_stern_test, 5 cases): diagonal operator
  against the closed-form complement solution, dense Hermitian
  U D U^dagger against the spectral reference with eps inside the
  occupied band, orthogonality of the solution to random occupied sets,
  degenerate and zero right-hand sides

Governance: the only findings are the two standing exemptions for this
design-phase module (header value-type includes <complex>/<vector>;
docs-sync with no user-visible INPUT change).

Verified: MODULE_DFPT_stern_test 5/5; ctest 9/9 (CELL 4 + DFPT 5);
abacus_pw_para links; governance --staged clean.

* Feat: first-order density response for DFPT (C3)

DFPT_Rho::compute_drho builds the q-shifted response density from the
Sternheimer solutions: the periodic parts u_nk (K-basis transform) and
du_nk (k+q coefficients scattered onto the rho grid through the shared
FFT-cell triple, C1 pattern) multiply pointwise into
A(r) = sum_{kn occ} wg u* du, real2recip gives the q-shifted coefficients
A_Delta = sum_{kn} wg sum_G c*_G d_{G+Delta} indexed by the rho-grid ig,
and the Delta = -q harmonic is dropped whenever -q falls on a reciprocal
lattice vector (charge conservation; always at q = Gamma). The manifest
real-space density 2 Re[e^{iqr} A(r)] is rebuilt from the projected
coefficients so both storages agree. mix_drho applies plain mixing on the
q-shifted coefficients through Base_Mixing::Plain_Mixing (zero initial
input, residual ||out-in||/||out||); the heavy Charge_Mixing header
dependency is replaced by a forward declaration plus a Matrix3 value
member (reciprocal matrix for q_frac -> cart).

- data layer: set/get_drho_r/set/get_drho_g go from stubs to real storage
- guards: nspin != 1 and non-plain mixing reject with WARNING_QUIT
  (design phase); cal_docc stays a documented U0 reservation (needs the
  PW-side beta-projector adapter wired with Plus_U in the C7/U1 window)
- unit tests (MODULE_DFPT_rho_serial, 5 cases): G-space coefficients
  against a brute-force double sum, real-space density against direct
  plane-wave sums, Gamma charge conservation, plain-mixing first/second
  step combination and residual formula
- test-side findings fixed (production code verified correct):
  PW_Basis_K::gcar is a per-k array indexed ik*npwk_max+igl
  (pw_basis_k.cpp:261) and must not be read with base-ball ig; direct-sum
  references must pair cartesian G with cartesian r = frac . latvec
- irrep wrapper test updated: drho storage slots are live (round-trip
  non-empty) after being design-phase stubs

Governance: only the two standing exemptions for this design-phase module
(value-type header includes, net dependency decreased by dropping
charge_mixing.h; docs-sync with no user-visible INPUT change).

Verified: MODULE_DFPT_rho_serial 5/5; ctest 10/10 (CELL 4 + DFPT 6);
abacus_pw_para links; governance --staged clean apart from exemptions.

* Feat: dynamical matrix for DFPT (C4 guard + C5)

- DFPT_Metal (C4): explicit WARNING_QUIT guards on the reserved metallic
  branch (dfdeps/compute_dmu/compute_drho_metal); interface-only as planned
- DFPT_Phon (C5):
  - ion_ion: Ewald force constants (G + R + self-image phase terms), the
    Gamma acoustic sum rule holds exactly by construction
  - accumulate_electron: 2n+1 complex accumulation 2 sum wg <dpsi^b|dV^a|psi>
    plus the same-atom anharmonic <psi|d2V|psi> term (d2vloc_r + apply_d2vnl
    from DFPT_Pert); the dpsi slot is backed up/restored around apply_dv
  - assemble/diagonalize/add_loto/check_sum_rule: zheev with signed cm^-1
    frequencies, LO-TO non-analytic term, Gamma row-sum rule
  - DFPT_PW_Data: dynmat stored as ComplexMatrix (complex Hermitian at
    generic q)
- fixes found by the new serial test: the cross term dropped the imaginary
  part (needed for the Hermitian symmetrization at q != 0) and the test
  reference used the basis momentum G instead of the kernel momentum G+q
- serial test MODULE_DFPT_phon_serial: 7 cases (Gamma ASR on a
  symmetry-broken two-atom cell, acoustic zero modes, incommensurate q vs
  direct dipole-Hessian sum, injected-dpsi closed-form contraction, zheev
  on a known matrix, isotropic LO-TO limit, Gamma sum rule)
- verification: 11/11 ctest targets pass (CELL 4 + DFPT 7), abacus_pw_para
  links, governance shows only the two pre-existing exempt warning classes

* Feat: q->0 response for DFPT (C6: eps, Born, v_hartree_q, XC contract)

- DFPT_Rho::v_hartree_q: q-shifted first-order Hartree kernel aligned with
  h_hartree_pw (skips |G+q|=0), shared by the C6 response and the C7
  screened potential
- XC_First_Order abstract contract in module_dfpt (adapter at the esolver
  layer in C7, mirroring DFPT_Stern::LinearOperator injection)
- DFPT_Pert::build_vkb_dk: analytic k-derivative of the beta projectors
  (atomic phase, radial chain rule, real-harmonic direction chain);
  build_vkb/build_vkb_dk made public for DFPT_Q0 reuse
- DFPT_Q0::pos_matrix: velocity (commutator) form
  r = -i <u|dH/dk|u> / (tpiba (eps_m - eps_n)), kinetic 2 tpiba^2 (k+G) plus
  the separable nonlocal derivative; degenerate pairs skipped
- DFPT_Q0::compute_eps / compute_born: length-gauge denominators, m sum
  over all bands for Z*, conj ordering of <v|dV|m>, ionic Z on the (a,b)
  diagonal, phon-style dpsi slot backup/restore
- serial tests: MODULE_DFPT_q0_serial (5 tests: vkb FD, kinetic analytic,
  nonlocal operator FD, eps two-level, born closed form) and v_hartree_q
  checks in MODULE_DFPT_rho_serial; 12-target regression + abacus_pw_para
  link pass

* Feat: wire DFPT driver and esolver factory (C7)

Module layer (C7a):
- DFPT_PW::init new signature (ucell, psi, bases, sf, veff_r, wg, eig,
  xc contract, nelec, ecutwfc, dftu); Impl holds GS data + hamilt_
- DFPT_HamiltShift: self-assembled H(k+q) Sternheimer operator
  (kinetic diagonal + veff FFT convolution + cached k+q vkb
  nonlocal), replacing the GS HamiltPW chain which is ik-index-bound
- DFPT_Pert::apply_vr public (screened response potential on all
  bands, FFT-cell triple core shared with real_space_dv)
- DFPT_Rho::reset_mixing per displacement; build_occ_kq folds k+q
  onto the GS k list; solve_displacement full SCF inner loop
  (v_hartree_q + xc_->apply -> RHS -> Sternheimer -> drho -> mix)
- run(): q=0 response + per-irrep displacement loop + assemble /
  diagonalize / add_loto; null-bases skeleton fallback kept

Esolver layer (C7b):
- ESolver_DFPT_PW: static config + inp-captured scalars in
  before_all_runners (rule 1: no global record re-read), run_gs ->
  init_dfpt wiring after SCF convergence (veff_smooth row, wg, ekb,
  psi, XC_First_Order_FDM adapter splitting Re/Im through PotXC_FDM)
- esolver.cpp factory 'dfpt' branch; read_inp_sys esolver_types
  + docs/parameters.yaml + input-main.md updated

Verified: ctest 12/12 (CELL 4 + DFPT 8); abacus_pw_para links;
-h esolver_type shows dfpt; --version v3.11.0-beta8. Governance:
1 allowed exception (determine_type factory PARAM read, existing
pattern) + known header/docs WARNINGs.

* Fix: DFPT screening-channel calibration (q=0 completion, XC central difference, per-displacement reset)

Three fixes verified against finite-difference references on the diamond
two-atom smoke case (optical 742.367x3 cm^-1 vs FD ~742, acoustic 6.40x3,
ASR residual 3.1e-6, off-irrep elements ~1e-11):

1. compute_drho: replace the in-place G-space Hermitian completion
   (double-processing each +-G pair, breaking Hermiticity and leaking a
   ~1.25x uniform overshoot) with a real-space 2 Re a(r) presymmetrization
   before real2recip; one-sided sticks whose -G falls outside the sphere
   now also complete correctly.
2. XC_First_Order_FDM: the forward difference Vxc[rho+drho]-Vxc[rho]
   carries a curvature term ~Vxc''*drho^2/2 that leaks a spurious A1
   component into v_sc (violating the A1xT2xA1 selection rule by 1.7e-2
   Ry/bohr) and destabilizes plain mixing at beta=0.7; use an eta=1e-6
   central difference instead (leak ~1e-11, default mixing converges).
3. solve_displacement: zero the stored drho_g when (re)entering a
   displacement so the previous response (or diverged leftovers) cannot
   leak into the first screening iteration.

Also includes the design-phase debug instrumentation used for the
diagnosis (DFPT_DEBUG/PTCHK/DYNCHK/MDBG/dump blocks, DFPT_MIX_BETA env
knob) and removes the VQCHK block that read PARAM.globalv.dq/nqx
(governance: keep the PR-level global dependency budget non-increasing).

Verification: ctest 10/10 (build/, MODULE_DFPT* + little_group + klist);
governance --staged clean except advisory warnings; smoke rerun after
VQCHK removal reproduces frequencies.

* Fix: DFPT plain-mixing default beta 0.7 -> 0.4 (small-G Coulomb stiffness)

The late-iteration divergence diagnosed in the diamond smoke case is a
plain-mixing stability issue, not a physics bug: residual stalls at 5e-5
then grows at exactly 1.2765x/iter while the iterate norm stays constant
(junk direction orthogonal to the physical component). The eigenmode is a
real Hermitian A1 breathing mode on the smallest G shells ({200} 6-vector
equal real amplitudes + {111} 8-vector +-pi/4 phases). A homogeneous
probe (inject the pure A1 trial, drop dV_ext from the rhs, measure the
one-iteration linear map; DFPT_JPROBE / DFPT_JPROBE_NOXC) gives

  lambda_A1 = -2.229   (Hartree-only -3.180, XC reduces it to -2.23)

i.e. the Coulomb stiffness 4pi/G^2 at small G. Plain mixing needs
beta < 2/(1+|lambda_min|) ~ 0.62; the physical T2 mode (lambda = -1.42,
less small-G head content) happened to converge at 0.7, which is why the
fixed point was correct while the A1 channel diverged (also explains the
earlier beta=0.3 convergence and the polluted drho manifest).

Default beta is now 0.4 (margin up to |lambda| ~ 5). Verification at
default settings: all six displacements exit via the convergence flag
(~38 iterations average, 228 total), frequencies identical to the
beta=0.7 forced run (optical 742.367 x3, acoustic 6.40 x3; fixed point
independent of beta), ele rows unchanged (e11 0.00286804 vs target
0.0028685, e12 -0.00286494 vs -0.0028701), converged drho manifest now
clean against the finite-difference reference (ratio 0.99994, cos
0.9993, 3.8% pointwise). ctest 10/10 (MODULE_DFPT* + little_group +
klist); governance --staged clean except advisory warnings. Proper fix
is a Kerker-type preconditioned mixer, noted for the B-phase follow-up.

Also adds env-gated design-phase diagnostics used for the diagnosis:
per-iteration residual print, MDBG dumps of drho/v_sc/v_ha/gcar, and the
JPROBE homogeneous-probe path.

* Docs: record DFPT stage-B gap audit and revised execution plan

* Feat: INPUT-driven DFPT parameters (dfpt_qmesh/qfile/compute_q0/loto/conv_thr/max_iter/mix_beta)

- read_inp_dfpt.cpp: 7 new INPUT items with checks (loto requires compute_q0)
- esolver_dfpt_pw: drop hardcoded qmesh/conv/max_iter and the dfpt.in stub;
  wire from inp explicitly (rule 1)
- DFPT_PW: set_qfile/set_mix_beta/set_compute_q0/set_loto; q file overrides
  the MP q mesh in init
- QList::read_from_file: fill the fallback A1 placeholder irrep (nirr=1)
  instead of clearing, so the q-file path keeps the 3N displacement fallback
- docs/parameters.yaml + input-main.md regenerated (new category)
- README example updated

* Test: sync DFPT serial references to production conventions

The four serial suites were last green against pre-calibration
binaries; three distinct reference gaps surfaced after the full
rebuild:

- pert/q0 AnalyticDVloc and FD references: the a004742 phase
  flip (GS stru_fac convention exp(-i 2pi g.tau), dVloc/dtau =
  -i (Delta+q)_alpha tpiba Vloc exp(-i 2pi (Delta+q).tau)) was
  not mirrored in the closed-form references.
- rho brute-force G-space and real-space manifests: compute_drho
  now carries the GS density normalization w/omega (elecstate
  rhoBandK w1); references divide by omega accordingly.
- phon accumulate_electron reference: same phase flip, plus the
  dynmat mass normalization /sqrt(m_a m_b) (term2) and /m (d2V)
  that the closed form had silently omitted (fixture mass 12).

MODULE_DFPT serial suites 26/26; full regression filter 14/14
(CELL 4 + DFPT 8 + IO 2). Governance: pre-existing exempted
include warnings only.

* Fix: multi-k DFPT ball-label matching and smeared-occupation projector cliff

Two independent defects broke DFPT responses whenever the ground-state
k list held more than one inequivalent point (nk > 1):

1. build_occ_kq assumed the k+q and k(q) balls share FFT-cell G labels.
   When k+q folds onto a different label of the same physical point
   (e.g. lists holding both (1/2,0,0) and (-1/2,0,0)), the projected
   states became garbage and the Sternheimer solve diverged. Balls are
   now matched through reciprocal-lattice integer triples
   f + dn = f', with dn = k(ik)+q-k(ikq); the ikq-side labels are read
   through PW_Basis_K::getgcar because collect_local_pw(erf) rebuilds
   gcar into a per-k ball layout [ik*npwk_max+igl], destroying the
   parent global-ig layout the old code indexed.

2. The absolute wg < 1e-8 occupied-band cliff made the Sternheimer
   projector jump between k samplings: a smeared Fermi-tail band with
   weight ~1e-6 sits on either side of the threshold depending on the
   sampling's Fermi level, opening or closing its empty-state channel
   in (H-eps)^-1 and shifting converged force constants by ~10%.
   A shared dfpt_band_occupied() now classifies a band as occupied
   iff wg(ik,ib) > 0.5*wg(ik,0) (majority occupation), applied
   consistently in the projector build, the solve driver, the response
   density, the 2n+1 assembly and the q0 valence/conduction split.

Diamond-Si 2-atom validation against finite differences (sym=0):
- single Gamma: D00 0.0208553 vs FD 0.020854 (unchanged)
- single L: D00 0.0129282 vs FD 0.012927 (new FD reference)
- {L,-L}: equals single-L exactly (was divergent), ASR row sums ~1e-6
- {Gamma,L}: D00 0.0166416 vs FD 0.016642 (was 0.0182462, +9.6%)
- {L,X} and weight-skewed {G,L} variants consistent; 14/14
  MODULE_DFPT/CELL/IO serial regressions pass.

* Docs: record multi-k DFPT root causes and FD validation matrix in PLAN

* Fix: reject metallic smearing occupations in DFPT with an explicit guard

An unshifted 2x2x2 mesh of diamond Si with the default gauss sigma
0.015 Ry places the smearing Fermi level 1.3 sigma below the Gamma VBM
(band occupations 0.92), and finite differences of the same ground
state then give force constants ~2.8x softer than DFPT: the E_f
response (d mu / d tau channel) is included automatically in any
finite-difference ground state but has no counterpart in the
Sternheimer flow (DFPT_Metal is a design-phase stub, C4). Without a
guard the run converges cleanly and reports silently wrong numbers.

DFPT_PW::init now scans the final wg and quits with an explicit
message when any band sits measurably between 0 and its full
reference (relative weight in (1e-3, 1-1e-3)); negligible gauss tails
are tolerated as the insulator limit.

Validation matrix for the regime boundary (diamond Si 2x2x2, sym=0):
- sigma 0.015: Gamma VBM 92% occupied -> guard fires (was 2.8x off FD)
- sigma 0.007: VBM 99.92% occupied -> guard passes, 3.8% off FD
  (residual dmu channel scales with tail weight)
- sigma 0.005: VBM 99.9996% occupied -> 0.05% off FD (insulator limit;
  D00 0.0127458 vs FD 0.012739), off-diagonals and ASR exact
Also validated in this round: single k=0.25,0,0 (D row0 real parts
match FD to 6e-7; imaginary antisymmetric parts are the expected
one-sided-k Hermitian artifact, the physical force constants are the
real parts), and single k=0.5,0,0 with symmetry=0 now reproduces the
L-point reference bitwise (symmetry=1 changes the single-k ground
state itself and is out of scope for FD comparison).

14/14 MODULE_DFPT/CELL/IO serial regressions pass. MPI>1 smoke
(-np 2) aborts with MPI_ERR_TRUNCATE in the DFPT phase: distributed
layouts are not yet supported and fail loudly.

* Docs: record validation-ladder extension, metallic-regime boundary, MPI smoke in PLAN

* Docs: record non-Gamma q smoke results (dfpt_qfile end-to-end, q<->-q consistency)

* Fix: drop spurious 1/nk in DFPT eps/born sums; wg already carries full BZ weight

compute_eps/compute_born divided the band sum by nk, but wg(ik,v) already
contains the full k weight wk times the spin factor 2, so the stored-k sum
is itself the BZ average. The extra 1/nk was a no-op for Gamma-only runs
(nk=1) and scaled down multi-k results by 1/nk.

Validation (Si diamond, LDA pz): 4x4x4 sym1 (8 IBZ k) eps_inf diagonal mean
= 12.6661; sym0 full-BZ 36 k manual sum = 12.6662 (5-digit cross-mesh
agreement; LDA reference ~12.7-13.2, experiment 11.7). Retained the
env-gated DFPT_Q0DBG p-matrix dump used for the parity-selection-rule audit.
Also documents in PLAN: wfc txt writer G-block (igl2isz FFT-stick order) vs
coefficient order (psi-ig) mismatch that invalidates file-based element-level
cross-checks, and the O_h parity selection-rule evidence that the in-code
p matrices are correct.

* Docs: record continuation plan (P0-1 uncommitted-fix intake, P0-2 Zstar bug, P0-3 B0 closeout, B2-B4, cleanup, A)

* Fix: gate same-atom d2V_ext on 2q reciprocal; drop spurious ion_ion delta/3

Physics (intake of the uncommitted 5-file fix, part 1 of 2):
- d2vloc_r: both displacement dressings e^{iqR} act on the same atom, so
  the cell sum collapses to G = 2q (mod ints); the local second-order
  kernel is nonzero only when 2q is reciprocal and then equals the plain
  q=0 integer-G kernel. Drop the dead q_cart parameter.
- apply_d2vnl: the second-order nonlocal operator carries wavevector 2q;
  build it on the q_eff = fold(2q)-shifted ball and gate the |dbeta><dbeta|
  middle projector term behind an explicit include_middle switch.
- accumulate_electron: apply the 2q-reciprocal gate to the whole d2 term
  (momentum-forbidden at generic q), pass q_eff/include_middle through,
  and fix the ion_ion same-image self term by removing the delta/3
  G=0 isotropic piece (validated element-wise against finite differences
  of the erfc-split Ewald energy in a q-commensurate supercell).
- ion_ion doc comment updated to the validated closed form.

Tests (dfpt_phon_serial):
- AccumulateElectronAnalyticContraction expectation synced to the
  Hermitian 2n+1 accumulation convention (commit dc82fac) and the
  gated-off d2 term at generic q; extract SetupBases(k, q) helper so a
  test can re-init the fixture at another (k, q).
- New AccumulateElectronD2GateOffGenericQ: row 0 stays pure cross at a
  generic q (gate suppresses the forbidden term).
- New AccumulateElectronD2CommensurateQ: k = (-1/2,0,0), q = (1/2,0,0)
  so 2q is reciprocal; three-component psi pins the cross term and the
  full d2 kernel K_{ab}(G) = -tpiba^2 G_a G_1 Vloc(G^2) e^{-i2pi G.tau}
  including the K(G_i - G_j) negative-harmonic convention.
- Zero the Psi buffers after construction (psi::Psi allocates
  uninitialized memory); without this the tests read heap garbage and
  become order-dependent in the shared-process serial suites.

* Debug: DFPT design-phase probes (ZDBG/BPT/NOSC/D2MID/DYNCHK/XB)

Part 2 of 2 of the uncommitted-fix intake: env-gated diagnostic probes
for the P0-2 Z* investigation and B-phase A/B debugging, all no-ops when
their env vars are unset (tracked for cleanup in
PLAN_dfpt_implementation.md probe ledger):
- DFPT_ZDBG (dfpt_q0 compute_born): per-occ-state decomposition of the
  Born-charge summand (wg, energy denominator, dV matrix element,
  position matrix element) to split occ-occ vs valence contributions.
- DFPT_BPT (dfpt_pw): perturbation-theory cross-check of the
  Sternheimer solve, <dpsi|rhs> vs sum_m |<psi_m(k+q)|rhs>|^2/(e_m-e_n)
  over the empty manifold at k+q (empty_kq_ cache added).
- DFPT_NOSC (dfpt_pw): zero the screened potential to isolate the bare
  Sternheimer response.
- DFPT_D2MID / DYNCHK d2gate (dfpt_phon): disable the |dbeta><dbeta|
  middle projector term; print the 2q-reciprocal gate decision.
- DFPT_XB: extend the row/column selection to the 2-atom rows 6.

Verified: MODULE_DFPT phon 9/9, q0 5/5 serial suites with probes inert.

* Docs: P0-1 done (2q-reciprocal d2 gate intake, order-dependence fix, 28/28 serial)

* DFPT q0: star-rotate the symmetry-reduced eps/Z* tensor sums

Symmetry-reduced k sums of the q=0 susceptibility tensors must be
star-averaged: the partial at a rotated star member Rk is R chi(k) R^T
(cartesian column form), with atom-resolved Born partials credited to
the image atom under the paired direct-space operation. The row-form
operator G^-1*kgmatrix*G from the kvec_d row convention had been fed to
rotate_tensor untransposed, which breaks the star sum (right- vs
left-coset representatives), so store the transpose.

Diamond Si 4x4x4 verification (sym=1): eps_inf = 12.6661*I and
Z* = 15.5799*I per atom, both bit-consistent with the symmetry-off
full-mesh reference (off-diagonals ~1e-14; previously 13.78/15.34/8.88
anisotropic). The remaining Z* offset vs the diamond target 0 is the
known missing-screening formula defect, tracked as the next P0-2 item.

Add StarRotationCyclicGroup to dfpt_q0_serial (C3 orbit cell: star size,
anisotropic trace-6 tensor averaging to 2*I, cyclic atom maps, identity
fallback) and the DFPT_STARDBG probe; build_stars/rotate_tensor/stars_
move to public for the test.

* DFPT q0: Sternheimer screened Z* (v4), QE-anchored eps 16pi fix, zstar_eu cross-check probe

- solve_pos_resp + compute_born v4: Y^a = (H-eps_v)^-1 P_c [H,x_a]|psi>
  (velocity rhs, build_vkb_dk nonlocal part), Z* = zion delta -
  2 sum wg Re <dpsi^kappa,scf|Y^a> (QE add_zstar_ue form); pos_resp/
  dpsi_efield stashes in DFPT_PW_Data
- eps factor 2 fix: 16 pi / Omega per QE dielec.f90 (8 pi was half);
  ComputeEpsTwoLevelAnalytic expectation synced, serial 6/6
- DFPT_ALEG probe: E-field SCF fixed point (solve_e form) + zstar_eu
  A-leg vs zstar_ue B-leg cross-check + SCF eps + DFPT_PTCROSS bare
  cross spectral diagnostic
- validated vs locally built QE 7.2 (same UPF/cell/ecut/mesh): GS energy
  identical, Gamma-TO 517.5/517.6 vs 517.63 (0.03%), Z* -1.19928 vs
  -1.19765 (0.14%), eps_scf 23.6825 vs 23.6685 (0.06%); 4x4x4 anomaly
  (Z*=-1.2, eps~23.7 vs lit 13) shown to be shared k-mesh convergence
  by QE discriminators (ONCV@4x4x4 same, pz-vbc@8x8x8 -> 14.04/-0.09)
- PLAN P0-2 closed with validation matrix and re-scoped acceptance

* DFPT q0: promote the E-field SCF solve, compute_eps to the dielec.f90 screened form

- solve_efield_resp is now production (QE solve_e order): runs after
  solve_pos_resp, before the displacement solves; converged dpsi^E,a
  stashed through DFPT_PW_Data (dpsi_efield)
- compute_eps consumes pos_resp + dpsi_efield:
  eps = 1 - (16 pi/Omega) sum_k wg sum_occ Re<Y^a|dpsi^E,b>, star-rotated
  on symmetry-reduced meshes; the PT r-matrix path is retired (pos_matrix
  kept as the design-phase analytic reference for its serial tests)
- serial test ComputeEpsScfSyntheticStash replaces the PT two-level case
  (prefactor, wg, occupied sum, conj/index pinning, empty-row skip); 6/6
- end-to-end sym 4x4x4: eps = 23.35 delta (was IPA 12.67), consistent
  with the nosym ALEG value 23.68 and QE dielec.f90 anchor 23.67

* DFPT: build_occ_kq diagnostic detail in the commensurability error; PLAN P0-3 intake (non-Gamma-q chain defect, eps SCF promotion record)

* DFPT: fix q!=Gamma phonon frequencies (missing spin factor 2 in drho), KQ dual-reservoir completeness, term3 d2 ungating

- compute_drho: include the spin factor 2 at every q (QE incdrhoscf wgt =
  2*weight/omega); the q=0 Hermitian completion now keeps Re only instead
  of 2 Re. Previously the screening was half strength away from Gamma,
  which collapsed the L-point Si frequencies to -948/-148/182/199 cm^-1.
  After the fix: 100.49/100.49/380.41/402.11/485.93/485.93 cm^-1 vs QE
  101.61x2/380.54/402.24/486.28x2 (Si NC 4x4x4, 0.1-1.1%); Gamma stays
  517.491 cm^-1 (QE 517.633).
- dfpt_kq_basis: dual-reservoir G assembly so the k and k+q balls share
  the same igl2ig maps (fixes silent truncation when one ball exhausts
  the rho-grid reservoir).
- dfpt_phon: drop the 2q-reciprocal gate on the same-atom d2 term (it is
  q-independent by construction; the old gate silently dropped it and
  produced imaginary branches).
- Verification: ctest 12/12 (MODULE_CELL x4 + MODULE_DFPT x8); serial
  4/4 (pert/phon/q0/rho); bare-response L run matches QE niter_ph=1 to
  0.008-0.4% (-2281.83 vs -2282.01 etc.).
- No docs change: module_dfpt is design-phase, no INPUT parameter touched.

* DFPT PLAN: P0-3 non-Gamma-q defect root-caused and fixed (drho spin factor 2, a915352)

* DFPT B2: formalize the phonon output (multi-q report, LO-TO corrected frequencies, data-layer loto direction)

- DFPT_PW_Data: loto_dir_ (unit-normalized setter, isotropic (1,1,1)/sqrt(3)
  default) and phon_freq_loto_ storage.
- DFPT_Phon: diagonalize_loto re-diagonalizes the Gamma matrix after
  add_loto and stores signed frequencies separately (plain phon_freq(0)
  stays intact); format_q_report/format_loto_report provide deterministic
  fixed-precision blocks (header with direct q coordinates and the
  correction direction).
- DFPT_PW::run uses data_.get_loto_dir() instead of the hardcoded
  (1,1,1)/sqrt(3); new accessors get_nq/get_qvec/get_loto_dir/
  get_phon_freq_loto/set_loto_dir plus the format forwarders.
- esolver run_post_process prints one block per q of the list plus the
  LO-TO Gamma block when enabled; tensor blocks only print when computed.
- Serial regression: 3 new cases (direction normalization, closed-form
  LO-TO spectrum {0, 13/12*pref}, char-exact format strings); phon 12/12,
  ctest 12/12, all 4 DFPT serial tests pass.
- End-to-end smoke (Gamma, compute_q0+loto, 4x4x4): TO 517.490709
  unchanged, LO-TO block along (0.577350 0.577350 0.577350), eps_inf
  23.6825 and Z*=-1.19928d for both atoms vs QE 23.6685/-1.19765 (0.13%).
  QE itself prints same-sign Z* with asr Sum=-2.395 for this setup; the
  acoustic-branch lift is the faithful consequence, not a defect.
- No docs change: module_dfpt is design-phase, no INPUT parameter touched.

* DFPT B3: Kerker-preconditioned density mixing in DFPT_Rho

- DFPT_Rho::init gains mix_type (plain/kerker) and kerker_a2 (1/lat0^2);
  no charge_mixing.h dependency, screen f_g = |G+q|^2/(|G+q|^2+a^2) built
  with the v_hartree_q convention (gcar + q_frac*G). Screen both inputs,
  plain_mix, add the screened part back: mixed = rin + beta*f*(out-rin)
  (QE semantics, stored density stays physical; |G+q|=0 harmonic frozen,
  consistent with its drop in compute_drho). Init signature extended with
  an explicit kerker_a2 argument (no default arg; both call sites updated).
- Wiring: env DFPT_MIX_TYPE / DFPT_KERKER_A2 design-phase knobs mirroring
  the DFPT_MIX_BETA precedent; default plain keeps behavior identical and
  the beta=0.4 default (and its stability rationale) stays documented in
  the init comment. No INPUT parameter change: no docs update required
  (env knobs are internal calibration aids, same category as DFPT_MIX_BETA).
- Tests (dfpt_rho_serial, 6 -> 8): analytic first Kerker step; lambda=-2.2
  stiff-shell model problem where plain beta=0.7 diverges (residual > 1)
  and kerker converges (< 1e-8) to the target.
- Fixed latent breaks masked by a stale test binary since a915352:
  kq0.init not updated to the 4-arg DFPT_KQ_Basis::init signature, and the
  brute-force references missing the band-weight spin factor 2.
- End-to-end (L point, 4x4x4, abacus_pw_para v3.11.0-beta8): plain beta=0.7
  diverges (|drho| -> 1e20); kerker beta=0.7 converges in 1393 s (vs 2332 s
  plain beta=0.4); frequencies identical across plain 0.4 / kerker 0.4 /
  kerker 0.7 to 8-9 digits (100.487828 x2 / 380.41385 / 402.10912 /
  485.93199 x2 cm^-1).
- Verification: OMP_NUM_THREADS=1 ctest -R 'MODULE_CELL_klist_test$|
  MODULE_CELL_reciprocal_grid_test|MODULE_CELL_qlist_test|
  MODULE_CELL_little_group_test|MODULE_DFPT' -> 12/12; serial suites
  pert 8 / phon 12 / q0 6 / rho 8 all pass; governance --staged clean
  except the expected no-docs-needed WARNING recorded here.

* DFPT B4: sink the (q,irrep) SCF ledger into DFPT_PW_Data, retire the DFPT_IrrepData adapter

- DFPT_PW_Data: the write-only single-slot ledger (set_current_iter(int)/
  set_converged(bool)/add_residual(double)) is replaced by the (q,irrep)-keyed
  six-accessor ledger sunk from DFPT_IrrepData (std::map value members,
  missing keys read as not-converged / empty history / iteration 0, clean()
  drops the ledger). The irrep dimension stays as the stage-A slot: the
  fallback irrep 0 carries the full 3N displacement basis. The new <map>/
  <utility> includes are required by the map value members the header owns.
- DFPT_IrrepData adapter deleted (git rm): its irrep==0 forwarding of
  dpsi/drho/dv duplicated the existing per-q data API, and its own keyed
  maps moved to the data layer. get_dpsi_obj (static dummy, zero callers)
  removed. Both CMakeLists updated, including the pw_run_test source list.
- run() outer-while accounting made honest: current_iter now increments per
  pass and convergence is worst-final-displacement-residual < conv_thr
  instead of an unconditional single pass. An unconverged pass re-runs the
  full solve (solve_displacement restarts from a zero input), bounded by
  max_iter outer passes, with the residual history keeping a record.
  Behavior on converged runs is bit-identical.
- solve_displacement / solve_efield_resp: write-only inner ledger writes
  removed; per-displacement state stays local and the final residual
  returns to run() for aggregation.
- Tests: dfpt_irrep_data_test.cpp renamed/rewritten as dfpt_pw_data_test.cpp
  (target MODULE_DFPT_pw_data_test, 5 cases: QList delegation, bound-safe
  accessors with the (q,spin) signature, setter round trip, keyed-ledger
  independence + clean() reset, U0 reservation).
- Verification: OMP_NUM_THREADS=1 ctest -R 'MODULE_CELL_klist_test$|
  MODULE_CELL_reciprocal_grid_test|MODULE_CELL_qlist_test|
  MODULE_CELL_little_group_test|MODULE_DFPT' -> 12/12 (pw_data_test fills
  the retired irrep_data_test slot); serial suites pert 8 / phon 12 /
  q0 6 / rho 8 all pass; end-to-end L-point default-config smoke
  (abacus_pw_para v3.11.0-beta8) reproduces the reference frequencies
  bit-consistently (100.487828/100.487829/380.413847/402.109158/
  485.931988/485.931988 cm^-1, TOTAL 2332 s, same as the pre-B4
  reference). Governance --staged: header-include warning justified by
  map value members; no INPUT behavior change so no docs update required.

* DFPT: retire the B-phase validation instrumentation (net -977 lines)

- Deleted (acceptance complete): PTCHK gauge/term2/HF-channel probes and the
  drho_dfpt.dat dump; the DYNCHK family (term2/d2gate/d2k/d2/ion/ele/elei and
  the DYNCHK4 double-zheev comparison); MDBG binary dumps (x2); JPROBE +
  JPROBE_NOXC (B3 acceptance done, delete as planned); OCCCHK incl. the
  dbg_miss label analysis and the empty_kq_/empty_kq_eig_ companion storage;
  XB; BPT incl. the want_empty projector expansion; NOSC; XCS/NOXC (v_sc
  assembly simplified to the knob-free path); DKCHK; YCHK; D2MID
  (include_middle sunk to literal true, q-independence settled); ALEG +
  PTCROSS (the whole aleg_crosscheck method); STARDBG; Q0DBG. Dead
  accumulators (d2sum_loc/nl, cross_k) and the now-purposeless <fstream>/
  <set> includes removed with them.
- Kept: DFPT_DEBUG (SCF residual tracing + posresp tracking, the B3/B4
  acceptance instrument and routine convergence diagnostics) and the B3
  calibration knobs DFPT_MIX_BETA / DFPT_MIX_TYPE / DFPT_KERKER_A2
  (documented in the DFPT_Rho::init comment).
- Behavior-preserving: every deleted probe was env-gated off by default;
  include_middle and want_empty defaults equal the sunk values.
- Verification: OMP_NUM_THREADS=1 ctest -R 'MODULE_CELL_klist_test$|
  MODULE_CELL_reciprocal_grid_test|MODULE_CELL_qlist_test|
  MODULE_CELL_little_group_test|MODULE_DFPT' -> 12/12; serial suites
  pert 8 / phon 12 / q0 6 / rho 8 pass; end-to-end L-point default-config
  smoke (abacus_pw_para v3.11.0-beta8) reproduces the reference frequencies
  bit-consistently (100.487828/100.487829/380.413847/402.109158/
  485.931988/485.931988 cm^-1). Governance --staged clean except the
  expected no-docs-needed WARNING (internal env probes, no INPUT change).

* delete PLAN

* Fix: adapt DFPT to the refactored Plus_U interface (compile break + U guard)

The develop-side DFT+U refactor (#7852-#7867) removed
source_lcao/module_dftu/dftu.h and the is_locale_initialized() member,
which broke every CMake build configuration of this branch at
dfpt_pw_data.cpp (all 9 CI build variants plus Test/CUDA/abacuslite
failed at the compile step; only the Makefile job passed because the
Makefile.Objects DFPT entries were absent at that merge point).

Changes:
- DFPT now consumes the PW-side Plus_U_Base (source_pw/module_pwdft/
  dftu_base.h) instead of the LCAO-side Plus_U header: dftu_ member,
  DFPT_PW_Data::init / DFPT_PW::init signatures and get_dftu() all use
  const Plus_U_Base* (the esolver call site passes &this->dftu with an
  implicit upcast). This removes the PW -> LCAO cross-layer include.
- u_active() = with_u() && is_occ_mat_initialized(): the reservation
  usability now follows the occupation-matrix state of the provider.
- DFPT_PW::init rejects a wired provider explicitly (WARNING_QUIT):
  the ground state supports PW-basis DFT+U now, but every DFPT U hook
  (cal_docc, build_dv_u, dftu_onsite, born/docc contractions) is a
  no-op U0 reservation, so running anyway would silently drop the
  whole first-order U response (fail-loud, same pattern as the
  metallic-sampling guard).
- test/dftu_test_support.cpp rewritten: the old static-member replicas
  no longer exist; the shim now provides only the Plus_U_Base ctor/dtor
  (also linked into MODULE_DFPT_pw_data_test, which constructs the
  provider directly). dfpt_pw_run_test's locale test becomes a death
  test pinning the WARNING_QUIT guard; the with_u/u_active contract
  moved to DFPT_PW_DataTest.DftuReservationProviderUsability; the
  unused dftu.h includes dropped from the phon/q0 serial tests.

Verification (GNU 8.3.1 + OpenMPI 5.0.3, GCC13 no-MPI cross-check):
- cmake --build build --target abacus_pw_para: builds/links
- cmake --build build-nompi (-DENABLE_MPI=OFF -DENABLE_LCAO=OFF)
  --target abacus_pw_omp: builds/links
- ctest -R 'MODULE_DFPT|MODULE_CELL': 50/50 pass (incl. the new death
  test and provider-usability case); ./build/abacus_pw_para --version
  prints v3.11.0-beta8
- agent_governance_check --staged: no findings

* Fix: reduce PR global dependency budget to non-increasing

The governance checker blocks the PR while the diff's added lines carry
more GlobalV/GlobalC/PARAM references than the removed lines
(added=51, removed=22, net_delta=+29 -> CI 'Governance checks' exit 1).
30 of the added references were test-side GlobalV::ofs_running streams;
they now use local std::ofstream objects (the fixture members that
already existed), and ReciprocalGrid::print_klists prints through its
own ofs parameter instead of the global stream (its single caller
passes the same running log). The stale 'Originally GlobalV::FINAL_SCF'
comment wording is dropped. The remaining production-side references
(reciprocal_grid.cpp k-point-file echo, klist.cpp MY_RANK guards) are
line-for-line moves of the previous klist.cpp code, so the budget is
now non-increasing (net_delta = -4).

Verification: ctest -R 'MODULE_DFPT|MODULE_CELL' 50/50 pass;
abacus_pw_para relinks; agent_governance_check --base origin/develop
--head HEAD exits 0 (no BLOCK findings).

* Fix: link K_Vectors/ReciprocalGrid sources into tests broken by the ReciprocalGrid refactor

The ReciprocalGrid refactor (Phase 1-3 of this PR) made K_Vectors
polymorphic: its vtable is now keyed on K_Vectors::renew and emitted in
klist.cpp, and the base vtable lives in reciprocal_grid.cpp. Twelve test
targets across estate/hsolver/stodft/io instantiate K_Vectors but never
compiled those translation units, so they fail to link after the merge
(masked until now by the earlier dftu.h compile break):

- MODULE_ESTATE_elecstate_{print,base,pw,energy}
- MODULE_PW_Sto_Hamilt_UTs
- MODULE_HSOLVER_pw
- MODULE_IO_write_bands (test_serial)
- MODULE_IO_write_eig_occ_test / write_dos_pw / print_info /
  read_wf2rho_pw_test (already had klist.cpp, lacked reciprocal_grid.cpp)
- MODULE_IO_write_dmk

Mirrors the pattern already used by this PR's own klist/qlist tests:
add klist.cpp + parallel_kpoints.cpp + k_vector_utils.cpp +
reciprocal_grid.cpp to SOURCES and the symmetry lib to LIBS.

Verified: full build green except MODULE_IO_numerical_basis_test (needs
ENABLE_LCAO, unguarded on develop as well); the fixed tests pass under
ctest; remaining local failures are environment artifacts (ScaLAPACK
abort-stub, ELPA off).

* Fix: link K_Vectors/ReciprocalGrid sources into LCAO-side tests and add new DFPT objects to Makefile.Objects

The ReciprocalGrid refactor made K_Vectors polymorphic (its key function
and the base vtable now live in klist.cpp / reciprocal_grid.cpp), so any
test that instantiates K_Vectors (module_dm tests, deltaspin
spin_constrain/template_helpers via spin_constrain.cpp, and
init_dm_from_file via density_matrix_io.cpp) fails to link.

Also register the five PR-added translation units (reciprocal_grid.cpp,
little_group.cpp, read_inp_dfpt.cpp, dfpt_hamilt_shift.cpp,
dfpt_kq_basis.cpp) in source/Makefile.Objects so the Intel Makefile build
does not fail with undefined references.

* Docs: resync parameters.yaml and input-main.md with the C++ Input_Item generator

The DFPT parameter block was hand-placed at a position that differs from
the item_dfpt() registration order, so the byte-exact consistency checks
in test.yml (--generate-parameters-yaml / generate_input_main.py) fail.
Regenerate both files with the documented commands to restore sync; the
only change is the position of the DFPT category block.

* Fix: compile reciprocal_grid.cpp in deepks unit tests

The ReciprocalGrid refactor made K_Vectors derive from
ModuleCell::ReciprocalGrid, so klist.cpp.o and k_vector_utils.cpp.o now
reference ReciprocalGrid member functions and its vtable. The
deepks_unit_support object library (DEEPKS_UNIT_COMMON_SOURCES, gated
behind ENABLE_MLALGO and thus only compiled in the gnu Test CI job)
compiles klist.cpp without reciprocal_grid.cpp, failing to link all 30
MODULE_LCAO_DEEPKS_* test executables with undefined references to
ModuleCell::ReciprocalGrid::renew/Monkhorst_Pack/build_star_ops/... and
its vtable/typeinfo. Add the missing translation unit to the common
source set; the symmetry library (incl. little_group.cpp) is already on
the link line.

* Fix: use threadsafe death tests in DFPT suites to avoid fork-in-threaded-process deadlock

MODULE_DFPT_pw_run_test timed out (1700 s) in the gnu Test CI job: the
two irrep-loop tests run first execute OpenMP regions, so with the job's
OMP_NUM_THREADS=2 the process is multithreaded when the third test
(dftu-reservation EXPECT_EXIT) forks. The default fast-style child then
deadlocks on exit and the parent waits forever (reproduced locally under
OMP_NUM_THREADS=2: gtest warns 'detected 2 threads' and hangs).

Switch all three DFPT death tests to the fork+exec threadsafe style
(same pattern as module_container tensor_test). For the pw_run test also
bridge std::cout to std::cerr inside the death statement: WARNING_QUIT
prints the NOTICE block to stdout, while death tests match the child's
stderr; the old CaptureStdout+HasSubstr assertion cannot see the re-exec
child's output. Verified under OMP_NUM_THREADS=2: pw_run 3/3 in 0.3 s
(previously indefinite hang), kq_basis 5/5, pert_serial 8/8, and the
full MODULE_DFPT ctest batch 8/8.

* Refactor DFPT unit tests: consolidate ctor/dtor stubs into shared dfpt_test_mocks.cpp (mirror tmp_mocks.cpp convention); absorb dftu_test_support.cpp

* Refactor DFPT unit tests: share the cubic-cell/stru_lib fixture between pw_data and pw_run tests (dfpt_stru_fixture)

* Refactor DFPT serial tests: derive pert/rho/phon/q0 fixtures from a shared DFPTSerialBase (cell/basis/data setup, Coulomb/NC atom builders, analytic dVloc reference)

* test(dfpt): dedupe repeated analytic blocks in the phon serial test

Share the occupied-weights table, the single-plane-wave psi builder, the
analytic accumulate_electron cross term (now on top of AnalyticDVloc),
and the isotropic loto data setup (eps/Born charges + two-atom mass
table via MakeTwoAtomCell) through phon fixture helpers; the three
AccumulateElectron tests and the two loto closed-form tests keep their
reference formulas but drop the duplicated inline copies.

* fix(dfpt): correct dn sign in k+q occupied-state ball folding

The #7894 refactor moved the k+q congruence matcher into
match_commensurate_kq but inverted its dn convention
(dn = k_d(ikq) - k_d(ik) - q) while the folding in
copy_occ_state_ball still used key = G + dn from the pre-refactor
convention (dn = k_d(ik) + q - k_d(ikq)).

Plane-wave identity requires G' + k_d(ikq) == G + k_d(ik) + q,
i.e. G' = G - dn, so for every k whose k+q folds across a BZ
boundary (dn != 0) the occupied-state coefficients were attached
to plane waves shifted by 2*dn. The Sternheimer operator then
developed ~+tpiba^2*|2dn|^2 errors in <psi|T+Vnl|psi> (~4 Ry for
Si@X, 4x4x4 mesh) and the CG solve diverged, yielding NaN phonon
frequencies for every non-Gamma q (48 of 64 k-points affected at
Si@X). The Gamma path (dn == 0 always) was unaffected.

Restores the pre-refactor mapping verified by the Si L-point
reference case.

---------

Co-authored-by: Zanthoxylum <chenshengjun@localhost.localdomain>
… hacks (#7989)

Five macros across three modules, each the last one in its module. The three
files are independent -- no shared production class -- and all follow the accessor
pattern established in #7984.

ions_move_methods_test (2 macros)
  Ions_Move_Methods already had public get_converged() and get_update_iter(), so
  every read was already covered; only the writes and the two aggregated
  sub-optimisers needed anything. It gains set_converged(), set_update_iter(),
  get_etot_info(), get_bfgs() and get_bfgs_trad(), and Ions_Move_BFGS2 gains a
  const get_is_initialized().
  The cross-object reads -- imm.bfgs.tr_min_hit, imm.bfgs.pos, imm.bfgs.inv_hess
  and the rest -- now chain through the Ions_Move_BFGS / BFGS_Basic accessors
  added in #7984, which is why that PR had to land first. Note a derived-fixture
  approach could not have worked here at all: C++ only lets a derived class reach
  a base's protected members through objects of its own type, and
  Ions_Move_Methods does not derive from BFGS_Basic.

esolver_dp_test (2 macros)
  runner() needs a real DP model file, so the test seeds the computed results and
  checks that cal_energy() / cal_force() / cal_stress() hand them back. Those are
  writes as well as reads, so ESolver_DP gains four reference accessors:
  get_atype(), get_dp_potential(), get_dp_force(), get_dp_virial().

lj_pot_test (1 macro)
  before_all_runners() derives the LJ tables in three steps and the test drives
  each on its own. ESolver_LJ gains six const accessors -- get_search_radius(),
  get_lj_rcut(), get_lj_c6(), get_lj_c12(), get_en_shift(), get_lj_virial() --
  and three wrappers: rcut_search_radius_for_testing(), set_c6_c12_for_testing()
  and cal_en_shift_for_testing(). All six reads are read-only here, hence const.

No production logic changed. No assertion or expected value changed.

Tree-wide count goes 28 -> 23, and source_relax, source_esolver and source_md
join source_base, source_basis, source_hamilt, source_hsolver, source_lcao,
source_main and source_psi at zero. What remains is source_io (11, being
restructured, so left alone), source_estate (10) and one file each in source_cell
and source_pw -- and #7987 / #7988 already take source_cell's and three of
source_estate's.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Refactor OFDFT line-search task buffer to std::array

* Use std::string throughout OFDFT line-search status handling

---------

Co-authored-by: zzlinpku <2601110378@stu.pku.edu.cn>
…to zero access hacks (#7987)

klist_test reached into K_Vectors through `#define private public`. It turned out
that only four of the things it touches actually needed anything, and two needed
nothing at all.

Needed wrappers -- set() is the single production entry point and drives these in
order, while the test exercises them one stage at a time because most of the
KPT-file parsing paths are only reachable that way:

  read_kpoints_for_testing()        27 call sites
  renew_for_testing()                8
  reduce_by_symmetry_for_testing()   2
  set_kup_and_kdw_for_testing()      5

Needed nothing:

  - spin_mult, 37 sites (36 writes, 1 read), already has public get_spin_mult()
    and set_spin_mult(); the setter was added in #7980 for write_dmk_test and
    covers every one of them.
  - koffset, 4 sites. Both Monkhorst_Pack() and read_kpoints() take the offset as
    an argument, and the 27 read_kpoints call sites already pass a local
    `const double koffset[3]`. Only two setup blocks wrote the member and then
    handed it straight back to Monkhorst_Pack(), so those now use a local array
    too and the member is not touched from the test at all.

The remaining 155 accesses in the file (get_nkstot, kvec_c, kvec_d, wk, kc_done,
kd_done, set_both_kvec, nmp, isk, ...) were public throughout and are unchanged.
I checked all six K_Vectors objects in the file -- kv, kv1 and kv_test1..4 -- and
the other classes the macro covered: nothing else private is used, and pseudo.h,
atom_spec.h, atom_pseudo.h and magnetism.h have no private sections at all.

Worth stating for review: the four wrappers are public entry points that only the
tests use, in the style of set_density_rotations_for_testing() in
symm_rotation.h, and carry a comment saying production code must keep going
through set(). read_kpoints() in particular is a substantial piece of behaviour
with 27 tests behind it; if review would rather see it simply become part of the
public interface, that is a one-line change and the wrapper can go.

No production logic changed. No assertion or expected value changed.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Mohan Chen <mohanchen@pku.edu.cn>
…ving elecstate_print_test's access hack (#7991)

All fourteen of elecstate_print.cpp's global reads were inside one function,
print_etot(), and it has exactly one production call site. print_etot() now takes
`const Input_para& inp` and the derived `two_fermi` flag, inserted before the
existing defaulted arguments so no new default argument is introduced.

elecstate_print.cpp is now PARAM-free. esolver_ks.cpp already holds an injected
inp_ and already passes *this->inp_ to ModuleIO::write_bands on the same path, so
the one call site follows a pattern that is already there.

The test owns an Input_para and a two_fermi bool instead of writing the private
half of PARAM. Its PARAM.sys.log_file write was dead -- no source compiled by
MODULE_ESTATE_elecstate_print reads it -- and is dropped. Passing the whole
Input_para rather than thirteen separate flags is deliberate: print_etot's job is
to report the INPUT-driven state, and thirteen parameters would be worse than the
aggregate it is actually printing.

No production logic changed. No assertion or expected value changed.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Fix(deltaspin): propagate current_spin through LCAO operator chain

HamiltLCAO::updateHk() sets current_spin on the root operator via
set_current_spin(isk[ik]), but the value was never forwarded to child
operators. DeltaSpin did not toggle its own current_spin, so it always
saw current_spin == 0 for nspin=2 and applied the same +lambda_z
coefficient to both spin channels instead of +lambda_z/-lambda_z. The
constraint therefore acted as a spin-independent potential and produced
wrong magnetic moments and total energies.

Propagate current_spin to the next operator in OperatorLCAO::init()
before processing the current node, so every node in the chain shares
the spin state set by the k-point loop.

Regenerate tests/03_NAO_multik/scf_deltaspin2/result.ref, whose previous
values encoded the buggy result.

* Fix(pw): correct nspin=2 DeltaSpin occupation output and per-atom labels

cal_occupations() read the becp layout with the npol=2 stride
(ib*2*nkb) and ignored the spin channel for nspin=2 (npol=1), so the
projected atomic magnetization printed by print_orb_chg() was wrong for
nspin=2. Index by the psi npol and store the spin-up/down occupancy in
the up-up/down-down Pauli blocks so that Charge = occ[0]+occ[3] and
Mag(z) = occ[0]-occ[3] print correctly; nspin=1 keeps a zero
magnetization; nspin=4 keeps the interleaved spinor layout.

Also build per-atom labels (Fe1, Fe2, ...) for the Total Magnetism /
Magnetic force tables in print_orb_chg(), print_Mi() and
print_Mag_Force(), instead of passing the per-type label vector (size
ntype) to tables with nat rows.

get_iat() is made const so the label helpers can read it from a
const SpinConstrain reference.

* Test: refresh stale DeltaSpin PW reference values

The result.ref of these five PW DeltaSpin cases predates the DeltaSpin
PW rework (they were last written in #7382) and no longer matched the
code: 12/18 were off by 2.6/4.6 eV, while 19/21/41 differed only at the
1e-4-1e-7 eV level. Regenerate all five with the current code so the
17_DS_DFTU suite passes again.

12_PW_DS_S2_Z now agrees with the dedicated 01_PW/scf_deltaspin2
reference (-6369.19826815 eV), confirming the new value is the intended
one.

* Refactor(pw): pass nspin explicitly to cal_occupations

The governance checker blocks PRs that increase GlobalV/GlobalC/PARAM
usage. cal_occupations() read PARAM.inp.nspin twice; pass it as an
explicit argument from ctrl_scf_pw() (which already holds the parsed
Input_para) instead. Also keep the print_orb_chg() header line using the
existing atom_label variable so the GlobalV::ofs_running line stays
untouched.

---------

Co-authored-by: dyzheng <zhengdy@bjaisi.com>
* module_charge: normalize indentation and brace single-statement control flow

Mechanical cleanup as the first step of the module_charge governance
refactor: convert leading tabs to 4-space indentation (1011 occurrences
across 11 files) and add braces around all single-statement if/for/while
bodies (11 sites). No functional change.

* module_charge: aggregate Charge_Mixing params into MixingConfig

Introduce a MixingConfig POD that bundles the INPUT mixing parameters
with the runtime globals (nspin, scf_thr_type, double_grid), and change
set_mixing from a 12-argument interface to set_mixing(const MixingConfig&,
double&, double&). Charge_Mixing now stores the config and reads nspin /
scf_thr_type / double_grid from it instead of PARAM.inp / PARAM.globalv,
removing the direct PARAM reads in set_mixing and init_mixing.

The single production call site (esolver_ks.cpp) fills the config, and
the unit test drives set_mixing via a make_cfg() helper. The
'#define private public' access hack is kept for now with a TODO: the
test still must write Parameter::input/sys, Charge::_space_* and
XC_Functional privates, which need the Step 4/5 global-state
parameterization before it can be removed.

Verified: make -j30 MODULE_ESTATE_charge_mixing (build_max_para_test)
passes with no errors.

* module_charge: deduplicate twobeta_mix lambdas and replace raw new with std::vector

Extract the repeated two-beta mixing functor in mix_rho_recip/mix_rho_real
into a make_twobeta_mix<T> template helper (6 lambda copies removed), and
convert all local raw new[]/delete[] buffers in charge_mixing_rho.cpp to
zero-initialized std::vector, dropping the paired ZEROS calls.

* module_charge: move residual/inner-product globals into MixingConfig

Extend MixingConfig with gamma_only_pw/domag/domag_z so mix_resid.cpp
(get_drho, get_dkin, inner_product_recip_{rho,simple,hartree,real}) no
longer reads PARAM/GlobalV; all branches now consume this->cfg_.
inner_product_recip_rho's raw pointer-array views are switched to
std::vector. Production fills the three new fields in esolver_ks, and
the test fixture gains a sync_cfg() helper to push PARAM mutations into
cfg_ for the inner-product branch tests.

* module_charge: own Charge's _space_* storage with std::vector (Step 5a)

Replace the six private raw _space_rho/_space_rho_save/_space_rhog/
_space_rhog_save/_space_kin_r/_space_kin_r_save buffers with
std::vector, so Charge's underlying contiguous storage self-manages and
the matching delete[] calls in destroy() (which relied on reading
possibly-uninitialized pointers) go away. The public rho/rhog/rho_save/
rhog_save/kin_r/kin_r_save views keep their double**/complex** shape and
still alias the vector memory via .data(), so all external consumers are
unaffected. Tests that drove _space_* directly are adapted to
resize()/.data() and drop their manual delete[] of the buffers.

* module_charge: route chgmixing_ks through its inp parameter

chgmixing_ks already takes a const Input_para& inp but still read
PARAM.inp.mixing_restart / PARAM.inp.scf_nmax from the global. Use the
inp argument instead so the function no longer reads INPUT state through
the global for these two fields. PARAM.globalv.ks_run is a runtime
per-process flag (set from band-parallel topology), not an input, so it
is intentionally left as-is rather than threading it through the
interface.

* module_charge: split Charge::init_rho into per-stage private methods

init_rho had a cyclomatic complexity of 36 from five sequential stages
(file read, atomic fallback, Thomas-Fermi tau, restart load, wfc read)
interleaved through shared read_error/read_kin_error flags. Extract the
four branches into private methods -- read_rho_from_file,
init_rho_atomic_and_tau, load_rho_from_restart, init_rho_from_wfc -- and
leave init_rho as a thin sequence of stage calls. Logic is unchanged; the
error flags are threaded through as parameters. The deepest stage
(read_rho_from_file) now sits at complexity 19, down from 36 for the
monolith. The remaining global reads inside the stages are untouched and
deferred to a later parameterization step.

* module_charge: extract Charge density math into charge_math free functions

sum_rho, cal_rho2ne and non_linear_core_correction each used Charge
members only to reach a handful of scalars (nrxx/nxyz/omega) or the
reciprocal-shell table (gg_uniq/ngg); the rest of each body is pure
numerics. Move the three bodies into a new charge_math namespace as free
functions with those values passed explicitly, and leave the Charge
members as thin forwarding wrappers so no caller outside the module
changes. The kernels are now unit-testable in isolation and no longer
coupled to Charge state. One behavior note: the pre-quit debug line that
printed sum_rho to ofs_warning is dropped so the free function stays free
of global-stream dependencies. charge_math.cpp is wired into the estate
library and the charge_test target.

* module_charge: register charge_math.o in the hand-written Makefile build

The CMake build already picks up charge_math.cpp; mirror that in
Makefile.Objects so the legacy Makefile flow links the new charge_math
kernels too. The module_charge directory is already on VPATH, so adding
charge_math.o to the object list is sufficient.

* module_charge: extract Charge::atomic_rho into charge_atomic free function

Remove Charge::atomic_rho entirely and replace all call sites with
module_charge::atomic_rho(..., rhopw), eliminating the need for a thin
wrapper on the Charge class. This decouples atomic density initialization
from Charge's state and improves charge.cpp quality score from 2 to 44.

* module_charge: forbid Charge copies and guard tau.cube write

scf_out_chg_tau aborted in Parallel_Grid::reduce on
assert(rhoin != nullptr) because the kin_r_save[is] handed to
write_vdata_palgrid was not a valid buffer. After the _space_* storage
became std::vector (ecf5084d4), a copied/moved Charge leaves its
rho/kin_r views dangling into another object's vector buffer, and a
kin_r_save never allocated (ked_flag set after allocate) stays nullptr;
both surface as a null rhoin deep inside MPI gather instead of at the
source.

Delete Charge's copy constructor/assignment so any value copy of the
vector-aliasing views fails at compile time, and check kin_r_save in
ctrl_output_fp before writing tau.cube so a missing allocation reports
a clear message instead of tripping the MPI assert.

Verification: not run locally (per user request, user compiles).

* module_base: tolerate null grid buffer when a rank owns no grid points

scf_out_chg_tau (LCAO, SCAN, out_chg=1, 4 MPI ranks) aborted in
Parallel_Grid::reduce on assert(rhoin != nullptr). Bisecting between
83eb5d0f3 (good) and ecf5084d4 (bad) isolated the regression to
ecf5084d4, which moved Charge's _space_* storage from raw new[] to
std::vector.

Root cause: with 4 ranks the FFT grid is slab-decomposed so that the
last rank owns zero real-space points (nrxx == 0, confirmed via a
temporary diagnostic printing fn/is/rank/nrxx at the reduce call site).
Before ecf5084d4, _space_rho = new double[nspin * 0] == new double[0]
returned a unique non-null pointer, so rho_save[is] was non-null and the
assert passed. After the change, an empty vector's .data() returns
nullptr, so the rank with nrxx == 0 handed a null rhoin to reduce and
tripped the assert (Debug) or fed MPI_Gatherv a null buffer (Release).

A rank with nrxx == 0 is legitimate: MPI_Gatherv is invoked with
sendcount 0 and ignores the send buffer. Relax the assert to only flag a
null buffer when nrxx != 0, and revert the now-unneeded kin_r_save guard
in ctrl_output_fp (it would have falsely aborted on the nrxx == 0 rank).

Verification: Release build (build_max_para_test), ran
  cd tests/03_NAO_multik/scf_out_chg_tau &&
  OMP_NUM_THREADS=1 mpirun -np 4 ../../../build_max_para_test/abacus_max_para
Result: exit 0, chg.cube and tau.cube written; numerical comparison
against chg.cube.ref/tau.cube.ref gives maxdiff 0 (chg) and 1e-14 (tau).

* module_charge: extract Charge::set_rho_core into charge_math free function

Move set_rho_core to charge_math::set_rho_core with rho_core,
rhog_core and rhopw passed explicitly instead of reading Charge
state, and call charge_math::non_linear_core_correction directly.
Remove the now-unused Charge::non_linear_core_correction wrapper,
use std::vector for the rhocg/vg scratch buffers, update the
init_scf call site, and drop the obsolete member stubs in the
elecstate unit tests.

* module_charge: vectorize Charge_Extra history arrays and forbid copies

Replace the raw new[]/delete[] displacement arrays (dis_old1, dis_old2,
dis_now) with std::vector and remove the hand-written destructor. This
fixes a read of uninitialized pot_order when an object is destroyed
before Init_CE, a memory leak when Init_CE is called repeatedly, and a
double-free risk from the implicitly generated shallow copy. The copy
constructor and copy assignment are deleted so the molecular-dynamics
trajectory history cannot be silently forked. The unit test now checks
vector sizes instead of non-null pointers.

* Rename charge_math to chg_tools and unify namespace module_charge

- Rename module_charge/charge_math.{h,cpp} to chg_tools.{h,cpp} via git mv
- Change namespace charge_math to module_charge to match charge_atomic
  and chgmixing in the same directory
- Update include guard CHG_TOOLS_H and TITLE/timer labels accordingly
- Update call sites in init_scf.cpp, charge.cpp, charge_init.cpp
- Update build references in Makefile.Objects and both CMakeLists.txt

* module_charge: refactor Symmetry_rho class to free functions

Convert the stateless class Symmetry_rho into namespace module_charge
free functions and rename files for consistency:
  symm_rho.{h,cpp}      -> chg_symm.{h,cpp}
  symm_rho_detail.h     -> chg_symm_detail.h
  symm_rhog.cpp         -> chg_symm_detail.cpp

- 5 public functions become module_charge::symmetrize_rho / cal_rhog_symm
  (2 overloads) / cal_rhog_symm_soc (2 overloads)
- 2 cross-TU helpers (psymmg/psymmg_soc) moved to module_charge::detail
  via chg_symm_detail.h
- 3 internal MPI helpers moved to anonymous namespace
- Delete dead code psymm (real-space symmetrization, never called)
- Remove empty ctor/dtor and parallel_grid.h include
- Rename begin/begin_soc to cal_rhog_symm/cal_rhog_symm_soc for clarity
- Update timer/TITLE labels from "Symmetry_rho" to "module_charge"
- Migrate all 14 call sites and 1 test stub
- Remove obsolete Makefile special rule (no more name collision)

* module_charge: extract MixingConfig header and drop unused inner_product_recip_simple

Move MixingConfig from charge_mixing.h into its own mixing_config.h so
stateless residual kernels can include the config without dragging in
Charge_Mixing. Remove inner_product_recip_simple, which had no production
call sites, together with its unit test.

* module_gint: move gint_prec_ctrl from module_charge

Relocate gint_prec_ctrl.{h,cpp} and its test into module_gint, update the
include in esolver_ks_lcao.h and rewire the CMake/Makefile object lists.

* module_charge: extract mixing inner products into chg_drho free functions

Rename mix_resid.cpp to chg_drho.cpp and turn inner_product_real and
inner_product_recip_hartree into module_charge free functions declared
in chg_drho.h; inner_product_recip_rho, which is only shared with the
unit test, moves to module_charge::detail in chg_drho_detail.h.
Charge_Mixing loses the three private inner-product members and
mix_rho_recip/mix_rho_real bind the free functions through lambdas.
get_drho/get_dkin stay as members for this step.

* module_charge: hide cal_drho/cal_dkin in an anonymous namespace

Move the get_drho/get_dkin implementations into file-local cal_drho/
cal_dkin free functions with all inputs explicit; the public
Charge_Mixing methods become thin forwarding wrappers so esolver call
sites stay unchanged.

* module_gint: fix include path in test_gint_prec_ctrl after relocation

* module_charge: extract Kerker screen kernels into chg_precond free functions

Move Charge_Mixing::Kerker_screen_recip/real to module_charge namespace
as free functions in chg_precond.{h,cpp}, renaming mix_precond.cpp via
git mv. Config/grid/geometry are passed explicitly via MixingConfig,
PW_Basis*, and tpiba, eliminating the function's direct read of
PARAM.inp.nspin. Replace 8 std::bind call sites in charge_mixing_rho.cpp
with lambdas, update 2 commented-out bind sites in charge_mixing_dmr.cpp,
and rewrite 12 test call sites in charge_mixing_test.cpp to construct an
independent MixingConfig instead of poking at Charge_Mixing privates.
Drop the now-unused member function declarations from charge_mixing.h.

* module_charge: fix Makefile.Objects after mix_precond -> chg_precond rename

Update the non-CMake object list to track the renamed translation unit so
make-based builds do not reference the deleted mix_precond.o.

* module_charge: drop Charge_Mixing::get_drho/get_dkin wrappers

Expose cal_drho/cal_dkin as module_charge free functions in chg_drho.h
and let ESolver_KS call them directly with explicit arguments; add
Charge_Mixing::get_mixing_config() as a const observer for the config.

* module_charge: rename chgmixing.h/cpp to chg_routine.h/cpp

Align with the chg_<feature> naming pattern used in the same directory
(chg_drho, chg_precond, chg_symm, chg_tools). Update include guard to
CHG_ROUTINE_H, the self-include in chg_routine.cpp, the entry in
source_estate/CMakeLists.txt and source/Makefile.Objects, and the three
#include sites in esolver_ks{,_pw,_lcao}.cpp. Function names
(chgmixing_ks{,_pw,_lcao}) and TITLE/timer tags are intentionally left
unchanged to keep the diff minimal.

* module_charge: rename mixing_config.h to chg_mix_cfg.h

Rename the MixingConfig header to align with the chg_* naming
convention in module_charge. Update the include guard and the four
in-tree includers; no CMake change is needed since the header is not
listed explicitly.

* module_charge: convert Charge MPI helpers into chg_parallel free functions

Rename charge_mpi.cpp to chg_parallel.cpp and add chg_parallel.h, moving
the three stateless Charge member functions (reduce_diff_pools, rho_mpi,
kin_r_mpi) to module_charge namespace free functions that take the
Charge object explicitly. Remove their declarations from charge.h and
update all call sites in elecstate_pw, stress_mgga, read_wf2rho_pw and
sto_iter. Rename the unit test to test_chg_parallel.cpp and update the
test target name accordingly.

GlobalV/PARAM reads and the direct MPI_Allreduce in reduce_diff_pools
are preserved as pre-existing technical debt (migration-neutral).

* Rename charge_atomic files to chg_atomic

- Rename module_charge/charge_atomic.{h,cpp} to chg_atomic.{h,cpp}
- Update include guard to CHG_ATOMIC_H
- Update includes in charge_init.cpp and charge_extra.cpp
- Update source paths in CMakeLists.txt, test CMakeLists.txt
- Fix stale object names in Makefile.Objects: replace
  symm_rho_charge.o/symm_rhog.o with chg_symm.o/chg_symm_detail.o

* module_charge: extract USPP double-grid split/merge into chg_uspp free functions

Introduce module_charge::split_dgrid / merge_dgrid in chg_uspp.{h,cpp} as
RAII, parameter-explicit replacements for Charge_Mixing::divide_data /
combine_data / clean_data, which paired raw new[] with manual delete[]
across ~160 lines of mixing code.

- chg_uspp.{h,cpp}: stateless free functions in module_charge namespace;
  outputs are caller-pre-sized std::vector, no new/delete; parameter
  validation via WARNING_QUIT; TITLE/timer tags preserved
- charge_mixing_rho.cpp: rho and tau double-grid paths switched to the new
  functions; raw pointer aliases kept for !double_grid so the existing
  mixing call sites (nspin==1/2/4) are untouched
- CMakeLists.txt (source + test): wire chg_uspp.cpp

The legacy divide_data/combine_data/clean_data members are not yet removed;
that follows in a later step after the test is updated.

* module_charge: rewrite MixDivCombTest for the new split_dgrid/merge_dgrid

Drop the legacy alias-pointer assertions (EXPECT_EQ(datas, data.data()),
EXPECT_EQ(datas, nullptr) after clean_data) that coupled the test to the
old new[]/delete[] ownership model.

The rewritten case verifies the actual contract:
- split_dgrid fills smooth and high-frequency buffers with the dense
  data verbatim (per-element comparison)
- merge_dgrid is a left-inverse of split_dgrid (output == input)
- no explicit cleanup call is required: std::vector manages storage

Covers nspin == 1 and nspin == 2 paths.

* module_charge: drop legacy divide_data/combine_data/clean_data members

With the new module_charge::split_dgrid/merge_dgrid in chg_uspp.{h,cpp}
and all call sites in charge_mixing_rho.cpp migrated, the original
Charge_Mixing::divide_data / combine_data / clean_data members are dead.

- delete charge_mixing_uspp.cpp (the raw new[]/delete[] implementation)
- drop the three member declarations from charge_mixing.h
- remove charge_mixing_uspp.cpp from source/test CMakeLists.txt
- Makefile.Objects: drop charge_mixing_uspp.o, add chg_uspp.o
- refresh one stale comment in charge_mixing_rho.cpp to reference
  merge_dgrid instead of the removed combine_data

* module_charge: rename charge_extra files to chg_extra and move class into namespace

Rename charge_extra.h/cpp to chg_extra.h/cpp and wrap the Charge_Extra
class in the module_charge namespace, matching the rest of module_charge
(chg_atomic, chg_symm, chg_uspp). Update include guards, call sites in
esolver_fp.h and the unit test, and CMake/Makefile source lists.

* module_charge: extract DMR mixing into chg_dmr free functions

Move the DMR allocation/mixing logic out of Charge_Mixing members into
stateless module_charge functions (init_mixing_dmr, template mix_dmr
with explicit instantiation), passing the Mixing object, mixing data
and MixingConfig explicitly instead of reading PARAM. Merge the two
identical real/complex mix_dmr overloads, replace raw new[]/delete[]
of the magnetic buffers with std::vector, and de-duplicate the
two-beta mixing lambda into a file-local helper. The members stay as
thin timer-wrapped wrappers so external call sites are unchanged.

* module_charge: remove Charge_Mixing DMR wrappers, call chg_dmr directly

Delete charge_mixing_dmr.cpp and have the two call sites
(chg_routine.cpp, esolver_ks_lcao.cpp) invoke module_charge::
init_mixing_dmr/mix_dmr directly with the Mixing object, mixing data
and MixingConfig obtained through Charge_Mixing accessors. Expose the
owned DMR mixing history via a new get_dmr_mdata() accessor and drop
the now-unneeded density_matrix.h include from charge_mixing.h.
Timers move into the free functions with module_charge labels.
Add the direct parallel_orbitals.h include to esolver_gets.h, whose
value member previously relied on the removed transitive include.

* module_charge: decouple chg_dmr kernel from HContainer, mix raw buffers

Change module_charge::mix_dmr to take per-spin raw contiguous double
buffers and nnr instead of HContainer/DMR container references, and
drop the hcontainer.h include (and its atom_pair/parallel_orbitals
dependency chain) from chg_dmr.cpp. The sole call site in
esolver_ks_lcao.cpp now extracts the wrappers and saved buffers from
the DensityMatrix containers before calling the kernel. Move the
argument checks into a file-local check_dmr_inputs helper. The kernel
now depends only on the mixing module and MixingConfig.

* module_charge: refactor charge_mixing_rho free functions and cleanup

- Replace 17 PARAM.inp/globalv direct reads with cfg_ fields
- Unify mixing_tau: remove redundant member, use cfg_.mixing_tau
- Extract make_twobeta_mix as free function template in anonymous namespace
- Extract mix_tau_recip free function for kinetic energy density mixing
- Extract pack_rho_mag/unpack_rho_mag templates for nspin==2 dedup
- Hoist screen and inner_product lambdas before if-else chains (8+4 dups)
- Remove dead new_e_iteration member and its no-op if block
- Drop unused parameter.h include from charge_mixing_rho.cpp

* module_charge: split member functions into charge_mixing.cpp, free functions into chg_rho_detail.h

- Move mix_rho_recip/mix_rho_real/mix_rho from charge_mixing_rho.cpp to charge_mixing.cpp
- Create chg_rho_detail.h for make_twobeta_mix, pack_rho_mag, unpack_rho_mag templates and mix_tau_recip declaration
- charge_mixing_rho.cpp now only contains mix_tau_recip definition in module_charge::detail
- Restore accidentally deleted mix_uom member function

* module_charge: rename charge_{init,mixing_rho} to chg_{init,tau}, widen cube_io ofs_running to ostream

* charge_init.{cpp,h} -> chg_init.{cpp,h}: move Charge::init_rho stages
  (read_rho_from_file, init_rho_atomic_and_tau, load_rho_from_restart,
  init_rho_from_wfc) from Charge member functions to module_charge free
  functions, dropping the corresponding private declarations from
  charge.h. Continues the module_charge convention of stateless free
  functions in chg_* files.

* charge_mixing_rho.cpp -> chg_tau.cpp: rename for the module_charge
  short-underscore convention; the file only contains mix_tau_recip.

* Extract mix_tau_recip declaration from chg_rho_detail.h into a new
  chg_tau.h so chg_tau.cpp no longer pulls in the detail template
  helpers (make_twobeta_mix / pack_rho_mag / unpack_rho_mag).
  charge_mixing.cpp adds chg_tau.h while keeping chg_rho_detail.h for
  the template helpers it still uses.

* Widen ModuleIO::read_vdata_palgrid's ofs_running parameter from
  std::ofstream& to std::ostream& (cube_io.h / read_cube.cpp). The
  body only uses operator<<, so std::ostream& is sufficient; this
  fixes the chg_init.cpp compile error where read_rho_file /
  read_kin_file (per project rules, std::ostream&) could not bind to
  the old std::ofstream& parameter. Existing callers passing
  std::ofstream& (GlobalV::ofs_running, test fixture) convert
  implicitly via base-class reference.

Build lists updated: source/Makefile.Objects and
source/source_estate/{CMakeLists.txt,test/CMakeLists.txt}.

Verification: chg_init.* changes compile-verified by user before
this session; chg_tau rename and chg_tau.h extraction not yet
compile-verified; cube_io type widening not yet compile-verified.

* module_charge: rename charge_mixing.{h,cpp} to chg_mix.{h,cpp}, test to test_chg_mix.cpp

Pure rename, no logic change. Updates include guard, 12 #include sites,
CMakeLists (source_estate + test), and Makefile.Objects. CMake target
MODULE_ESTATE_charge_mixing kept (no external references). Class name
Charge_Mixing and module_charge namespace unchanged.

* module_charge: remove duplicate doc block comments (Phase 1a)

Remove or rephrase 14 duplicate comment lines across 7 files to
eliminate all duplicate_doc_block quality-score deductions.

- chg_mix.cpp: remove 7 duplicate comments in mix_rho_real that
  repeated mix_rho_recip's broyden/Kerker/magabs annotations
- chg_init.cpp: remove 2 duplicate comments in read_kin_file that
  repeated read_rho_file's binary-read and ParaWorld bridge notes
- chg_symm_detail.cpp: remove 1 duplicate step comment in psymmg_soc
- charge.h: rephrase kin_r_save comment to avoid repetition
- chg_extra.h: rephrase beta comment to avoid repetition
- chg_symm.cpp: remove 1 duplicate vector-management comment
- chg_precond.cpp: remove 1 duplicate Kerker comment

* module_charge: replace auto with explicit std::function types (Phase 1b)

Replace 14 auto-keyword lambda declarations with explicit
std::function types to eliminate all auto_keyword quality-score
deductions.

- chg_mix.cpp: 10 auto -> std::function (inner_product, screen,
  twobeta_mix in mix_rho_recip and mix_rho_real)
- chg_drho.cpp: 2 auto -> std::function<double()> (part_of_noncolin,
  part_of_rho)
- chg_tools.cpp: 1 auto -> std::function<void(int,int)> (kernel)
- chg_symm_detail.cpp: 1 auto -> std::function (build_wspin)

Added #include <functional> to all four files.

* module_charge: wrap lines over 120 chars (Phase 1c)

Break 21 lines exceeding the 120-char limit across 7 files to
eliminate all line_too_long quality-score deductions.

- charge.cpp: 3 WARNING_QUIT/cout lines split
- chg_atomic.cpp: 5 Simpson_Integral/exp/assert lines split
- chg_drho.cpp: 2 conj-product sum lines split
- chg_init.cpp: 1 warning message string split
- chg_mix.cpp: 5 make_twobeta_mix/recip_to_real/if_scf_oscillate lines split
- chg_mix.h: 3 member declaration/comment lines shortened
- chg_symm_detail.cpp: 2 MPI_Recv lines split

* module_charge: remove default parameter from Charge::init_rho (Phase 1d)

Remove the default nullptr values from init_rho's klist and wfcpw
parameters and update the two call sites (esolver_of.cpp,
esolver_double_xc.cpp) that relied on the defaults to pass nullptr
explicitly.

* module_charge: replace raw new/delete with std::vector and unique_ptr (Phase 2a-2d)

Replace all raw new/delete allocations in 4 files with RAII
containers to eliminate raw_new_keyword and unpaired_new_delete
quality-score deductions.

- chg_tools.cpp: 1 new -> std::vector<double> (aux buffer)
- chg_extra.cpp: 4 new -> std::vector<std::vector<double>> (rho_atom
  in extrapolate_charge and find_alpha_and_beta)
- chg_symm_detail.cpp: 14 new -> std::vector (rhog_piece, ig2isz,
  ipsz2ipw, nstnz_start, fftixy2is, rhogtot, ig2isztot, ixyz2ipw
  across reduce_to_fullrhog, rhog_piece_to_all, psymmg, psymmg_soc)
- chg_mix.{h,cpp}: 5 new + 5 unpaired -> std::unique_ptr for
  mixing and mixing_highf members; destructor and init_mixing
  simplified; get_mixing() returns .get()

charge.cpp (18 raw new) deferred to Phase 2e due to wider impact.

* module_charge: replace raw new/delete in Charge with vector-backed storage (Phase 2e)

Replace all 18 raw new and 10 unpaired delete in charge.cpp with
std::vector-backed storage to eliminate raw_new_keyword and
unpaired_new_delete deductions.

- charge.h: add _ptrs_rho, _ptrs_rhog, _ptrs_rho_save, _ptrs_rhog_save,
  _ptrs_kin_r, _ptrs_kin_r_save (std::vector<double*> / complex*),
  and _space_rho_core, _space_rhog_core (std::vector data buffers)
- charge.cpp allocate(): replace new double*[nspin] with vector resize;
  rho = _ptrs_rho.data() preserves double** interface
- charge.cpp init_final_scf(): replace both outer pointer and inner
  data new calls with _space_* vectors
- charge.cpp destroy(): replace delete[] with vector::clear() and
  nullptr assignment

charge.cpp score: 47 -> 69, now passing the 60 threshold.
Module average: 85.0 -> 85.7, 30/33 files passing.

* module_charge: replace std::make_unique with C++11-compatible unique_ptr(new T) (fix)

std::make_unique is a C++14 feature; the repo baseline is C++11.
Replace 4 make_unique calls with std::unique_ptr<T>(new T(...)) to
eliminate the post_cpp11_feature deduction (-40).

chg_mix.cpp score: 0 -> 15, module average: 85.7 -> 86.1.

* module_charge: fix duplicate doc block in charge.cpp init_final_scf

* module_charge: aggregate chgmixing_ks parameters into ScfMixingCtx struct (Phase 3a)

Replace 14-parameter chgmixing_ks with 7-parameter version by
grouping SCF convergence thresholds and status flags into a new
ScfMixingCtx struct, and deriving nrxx from chr.rhopw->nrxx.

- chg_routine.h: define ScfMixingCtx struct (hsolver_error, scf_thr,
  scf_ene_thr, converged_u, drho, oscillate_esolver, conv_esolver)
- chg_routine.cpp: unpack ctx members at function entry
- esolver_ks.cpp: pack ctx before call, unpack after

chg_routine.cpp score: 63 -> 70, too_many_parameters eliminated.

* module_charge: aggregate read_rho_file/read_kin_file parameters into ReadCfg (Phase 3b)

Replace 9-parameter read_rho_file and read_kin_file with 5-parameter
versions by grouping suffix, readin_dir, rank, ofs_running, ofs_warning
into a ReadCfg struct in the anonymous namespace.

chg_init.cpp score: 66 -> 70, too_many_parameters eliminated.

* module_charge: aggregate non_linear_core_correction parameters into NlcCtx (Phase 3c)

Replace 10-parameter non_linear_core_correction with 2-parameter
version by grouping all input data into a new NlcCtx struct.

chg_tools.cpp score: 96 -> 100, too_many_parameters eliminated.

* module_charge: split chg_mix.cpp into init and rho mixing files (Phase 4a)

Move mix_rho_recip, mix_rho_real, and mix_rho (440 lines) from
chg_mix.cpp into a new chg_mix_rho.cpp to eliminate file_too_long
deduction (-10).

- chg_mix.cpp: 727 -> 286 lines (constructor, set_mixing,
  init_mixing, set_rhopw, mix_reset, if_scf_oscillate,
  allocate_mixing_uom, mix_uom)
- chg_mix_rho.cpp: new file, 440 lines (mix_rho_recip,
  mix_rho_real, mix_rho)
- CMakeLists.txt: add chg_mix_rho.cpp to library and test targets

chg_mix.cpp score: 15 -> 60, now passing the 60 threshold.
32/34 files passing, module average improved.

* module_charge: split chg_drho.cpp and decompose inner product functions (Phase 4b)

Move inner_product_recip_rho and inner_product_recip_hartree from
chg_drho.cpp into a new chg_drho_inner.cpp, and decompose each
into per-nspin helper functions to reduce cyclomatic complexity.

- chg_drho.cpp: 520 -> 161 lines (cal_drho, cal_dkin,
  inner_product_real); score 49 -> 97
- chg_drho_inner.cpp: new file, 310 lines; score 100
  - inner_product_recip_rho decomposed into recip_rho_nspin1,
    recip_rho_nspin2, recip_rho_nspin4_mag helpers (CC 29 -> ~5 each)
  - inner_product_recip_hartree decomposed into
    recip_hartree_nspin2, recip_hartree_nspin4_trad,
    recip_hartree_nspin4_angle helpers (CC 37 -> ~5 each)
  - shared coulomb_sum_single extracted
- CMakeLists.txt: add chg_drho_inner.cpp to library and test targets

34/35 files passing, only chg_atomic.cpp remains below 60.

* refactor(module_charge): split atomic_rho and remove ZEROS in charge mixing

chg_atomic.cpp:
- Decompose atomic_rho (CC=60) into per-nspin helpers in
  chg_atomic_inner.cpp; CC reduced to 7, score 40->100.
- Replace all PARAM.inp.nelec/domag/domag_z/test_charge and
  GlobalV::ofs_warning with explicit AtomicRhoCfg parameter.
- Remove unused parameter.h include.
- Add chg_atomic_detail.h declaring detail helpers and RhoG3dCtx.

chg_init/chg_extra/esolver_*:
- Pass AtomicRhoCfg through call sites of atomic_rho,
  extrapolate_charge, and update_delta_rho.

Bug fixes:
- chg_drho_inner.cpp: fix duplicate const (const MixingConfig const&
  -> const MixingConfig&) and add detail:: prefix to helper calls.
- chg_mix_rho.cpp: use mixing.get()/mixing_highf.get() for unique_ptr.
- chg_tools.cpp: fix numeric -> numeric[it] in set_rho_core.

Memory safety / cleanup:
- Replace ModuleBase::GlobalFunc::ZEROS with std::fill in charge.cpp,
  chg_symm_detail.cpp, chg_tools.cpp; remove redundant ZEROS calls
  that precede full overwrites in chg_dmr.cpp and chg_mix_rho.cpp.

* Refactor: remove redundant Charge& overload of cal_rhog_symm_soc

The Charge& overload only forwarded chr.rho/chr.rhog to the raw-array
overload and had a single internal call site. Inline the member access
at that call site and drop the wrapper declaration and definition.

* module_charge: fix stale TITLE/timer labels and drop unused xc_functional.h includes

mix_tau_recip is now a free function in module_charge::detail, so update
its TITLE/timer labels from the legacy "Charge_Mixing" to "module_charge"
to match the convention of other free functions in the directory. Also
remove the unused xc_functional.h includes from chg_tau.cpp and
chg_symm_detail.cpp (label/include cleanup only, no behavior change).

* module_charge: remove redundant #ifdef __MPI guards around parallel wrappers

Parallel_Reduce::reduce_pool and Parallel_Common::bcast_double already
compile to no-op stubs when __MPI is undefined, so the outer guards add
nothing. Remove 11 such guards in chg_tools.cpp, chg_drho.cpp,
chg_drho_inner.cpp, chg_atomic_inner.cpp and chg_mix.cpp.

Guards enclosing raw MPI calls or MPI/serial dual paths are kept
(chg_parallel, chg_symm_detail, chg_routine BP_WORLD bcast, chg_extra.h).

* module_charge: decouple chg_routine from spin_constrain singleton

- forward-declare Plus_U_Base in chg_routine.h instead of including dftu_base.h
- query DeltaSpin mag_converged in ESolver_KS_PW and pass it to chgmixing_ks_pw

* module_charge: remove PARAM dependencies via explicit configuration structs

Remove the last four direct includes of parameter.h in module_charge
(chg_mix, chg_parallel, charge, chg_init) and the implicit PARAM.globalv.ks_run
read in chg_routine. INPUT values are now passed explicitly:

- MixingConfig gains scf_nmax for the drho oscillation history
- reduce_diff_pools/rho_mpi/kin_r_mpi take kpar, all_ks_run, bndpar, nspin,
  out_elf from callers instead of GlobalV::KPAR/PARAM
- Charge::kin_density/allocate/check_rho/renormalize_rho/init_final_scf take
  out_elf/test_charge/nelec as arguments with validation asserts
- new InitRhoCfg aggregates INPUT values for init_rho
- ScfMixingCtx gains ks_run; dm2rho takes nelec and drops its default
  skip_normalize argument per governance rule 5

No behavior change: save_rho_before_sum_band now uses the member nspin
set by allocate, identical to the previously read PARAM.inp.nspin.

* module_charge: restore #ifdef __MPI guards around parallel wrapper calls

The guards removed in 7a0013848 are load-bearing for serial-built unit
tests: source_estate/test strips __MPI from test translation units via
abacus_disable_feature_definitions, but links libbase built with __MPI,
whose explicit Parallel_Reduce instantiations contain real MPI calls.
Unguarded calls in the test TUs therefore bound to MPI_Allreduce and
abort with "called before MPI_INIT", failing MODULE_ESTATE_charge_test
and MODULE_ESTATE_charge_mixing.

Restore all 11 call-site guards in chg_tools.cpp, chg_atomic_inner.cpp,
chg_drho.cpp, chg_drho_inner.cpp and chg_mix.cpp. No behavior change for
MPI or serial production builds.

* Remove dead PAW compensation charge members

nhat, nhat_save in Charge and nhat_mdata in Charge_Mixing have had
no references since #6225 removed the PAW code; drop the orphaned
declarations and update the related comment.

* Refactor: remove unused Charge::prenspin member

prenspin recorded the spin-channel count read from legacy cube charge
files and drove collinear-to-noncollinear rearrangement in init_rho.
After read_rho was replaced by binary read_rhog (#5323, #5362) the
value is neither written nor read anywhere, so drop the dead member.

* Refactor: move Charge::cal_rho2ne/check_rho to module_charge free functions

- Add module_charge::check_rho in chg_tools.{h,cpp} with grid/geometry
  parameters passed explicitly; preserve all branches, thresholds and
  warning/abort messages of Charge::check_rho
- Remove the Charge::cal_rho2ne forwarding wrapper and Charge::check_rho
- Update the three esolver call sites (ks/of/double_xc) to pass rho,
  nspin, rhopw grid sizes and ucell.omega explicitly
- Drop the check_rho stubs in elecstate_pw/base tests and switch
  charge_test to the free cal_rho2ne
- Add test_chg_tools.cpp covering cal_rho2ne, total/spin-polarized
  checks, mismatch warning path and negative-channel aborts

* Refactor: remove redundant Charge::omega_ pointer

- Charge::sum_rho() now reads the cell volume from rhopw->omega, which
  is computed from the same lat0/latvec as ucell.omega and is already
  dereferenced on the same line for nxyz; this also makes the volume
  consistent with the grid rho lives on
- Drop the Charge::omega_ member, its set_omega() setter and the
  chg_init.cpp call site, removing a raw-pointer dependency on the
  UnitCell lifetime; update charge_test accordingly

Verified: MODULE_ESTATE_charge_test and MODULE_ESTATE_chg_tools pass,
elecstate library rebuilds cleanly.

* Remove dead Charge::init_final_scf and allocate_rho_final_scf

init_final_scf has had no production callers since the nscf refactor
(c6ae01236); its only remaining caller was the unit test added in
ba8b7ce9a. After the vector-backed storage refactor it was also a
broken duplicate of Charge::allocate: it never set nspin/nrxx/nxyz/
ngmc and skipped the kin_r buffers. Remove the function, its one-shot
guard flag, and the corresponding test case; destroy() now keys solely
on allocate_rho since vector storage self-manages cleanup.

* Refactor: pass rhopw explicitly to chg_init/chg_routine/chg_extra/chg_symm

Remove implicit reads of chr.rhopw/chr.ngmc from four module_charge files:
- chg_symm.cpp: size kin_g by the rho_basis used for its FFTs
- chg_routine: chgmixing_ks takes const PW_Basis&
- chg_init: orchestrator and four stage helpers take const PW_Basis&;
  the Charge::init_rho member signature is unchanged
- chg_extra: extrapolate_charge/update_delta_rho take const PW_Basis&

Call sites pass *chr.rhopw at the KS boundary or *pw_rhod where the
binding (esolver_fp.cpp chr.set_rhopw(pw_rhod)) makes them identical.
Verified: affected TUs compile and MODULE_ESTATE_charge_extra passes.

* Comments: add TODOs for LCAO+USPP double-grid follow-ups

Record the smooth/dense grid split to revisit if LCAO is ever allowed
with USPP: symmetrize_rho callers pass different grids, and the
ndx/ndy/ndz input path lacks the LCAO guard the ecutrho path has.

* Refactor: replace sticky Charge::cal_elf flag with explicit symm_kin argument

cal_elf was set to true once during ELF output and never reset, so every
later density symmetrization in the same run redundantly symmetrized
kin_r. Replace the mutable workflow flag with an explicit bool parameter
on the Charge& overload of module_charge::cal_rhog_symm:
- ctrl_output_fp passes true right before write_elf consumes kin_r
- symmetrize_rho wrapper and other callers pass XC_Functional::get_ked_flag()

Verified: full incremental build, read_wf2rho unit tests (serial/4 MPI),
write_elf logic test, and tests/01_PW/scf_out_elf (E difference 5e-10 eV,
ELF cube passes CompareFile.py at 3-decimal tolerance).

* Refactor: resolve mixing_tau at config assembly, drop XC dependency from chg_mix

esolver_ks now resolves mix_cfg.mixing_tau = inp.mixing_tau &&
XC_Functional::get_ked_flag() at the single production config assembly
point, so chg_mix/chg_mix_rho no longer query the XC global inside tau
mixing branches (6 sites). test_chg_mix mirrors the resolution in
make_cfg() and sets ked_flag before set_mixing where tau mixing is
expected. Also drop an unused xc_functional.h include from
chg_drho_inner.cpp.

Verified: full incremental build clean; MODULE_ESTATE_charge_mixing
11/11 tests pass; MODULE_ESTATE_charge/chg test suites all pass
(serial + 4-rank MPI).

* Fix: restore complete types in chg_drho_inner.cpp after include removal

Removing xc_functional.h in 87b818f4c broke compilation: the include was
load-bearing transitively, supplying the complete ModulePW::PW_Basis type
and ModuleBase::TITLE. Add the direct includes instead (pw_basis.h,
global_function.h) per IWYU.

Verified: make -j16 exits 0 with full log retained (previous verification
was invalid: a tail pipe masked both the exit code and the errors).

* Refactor: derive tau symmetrization/reduction from kin_r buffer existence

The Charge& cal_rhog_symm overload and rho_mpi/kin_r_mpi queried
XC_Functional::get_ked_flag() (plus a caller-supplied out_elf/symm_kin
flag) to decide whether to touch kin_r. Since Charge::allocate allocates
kin_r exactly when meta-GGA or ELF output needs it, both now check
chr.kin_r != nullptr directly, dropping the XC dependency and the extra
boolean parameters:
- rho_mpi/kin_r_mpi lose the out_elf parameter (2 production, 3 test
  call sites updated)
- the Charge& cal_rhog_symm overload loses the symm_kin parameter
  (ctrl_output_fp, setup_pot, read_wf2rho, update_state_rdmft revert to
  4 arguments); the raw-pointer overload now checks kin_r != nullptr
  only
- module_charge keeps XC references only in charge.cpp, chg_init.cpp,
  chg_drho.cpp (semantic "is meta-GGA" sites, resolved next)

Verified: make -j16 exit 0; 14/14 ctest charge/elecstate/read_wf2rho
tests (serial + 4-rank MPI); tests/01_PW/scf_out_elf integration case
reproduces the reference energy (-194.623411265 eV, diff 5e-10) and the
ELF cube passes CompareFile.py at 3-decimal tolerance.

* Refactor: remove module_xc dependency from module_charge (meta_gga state)

module_charge queried XC_Functional::get_ked_flag() at 5 semantic
"is meta-GGA" sites (tau TF init, tau file read, tau save, tau residual,
tau mixing resolution). Resolve the flag at upper layers instead:
- Charge::allocate takes an explicit meta_gga argument and stores it as
  object state; save_rho_before_sum_band and cal_dkin read it
- InitRhoCfg gains a meta_gga field, filled at the 3 esolver config
  assembly points (ks/of/double_xc)
- delete Charge::kin_density(); 6 esolver call sites inline
  get_ked_flag() || (out_elf[0] > 0) for buffer allocation and pass
  get_ked_flag() as meta_gga; non-SCF allocations pass false
- charge_test mirrors the inline expression

module_charge now has zero references to module_xc.

Verified: make -j16 exit 0 (full log); 14/14 charge/elecstate/
read_wf2rho ctests (serial + 4-rank MPI), including the mGGA tau mixing
and tau-save branches; tests/01_PW/scf_out_elf reproduces reference
energy (-194.623411265 eV, diff 5e-10) and the ELF cube passes
CompareFile.py at 3-decimal tolerance. A SCAN integration case
(205_PW_SCAN) still requires a libxc-enabled build/CI run.

* Fix: allow null rho buffers on ranks with empty real-space grid partition

pack_rho_mag/unpack_rho_mag in chg_rho_detail.h quit whenever any buffer
pointer is null. A rank may legitimately own zero real-space grid points
(nrxx == 0) when the grid is decomposed across more processes than it has
z-slabs (e.g. a 3x3x3 big-cell grid on 4 processes leaves one rank with
no slab); its zero-sized vectors then return null data() pointers even
though the packing loops perform no access. The unconditional check made
LCAO nspin==2 real-space mixing abort with "pack_rho_mag pointer is null"
on such ranks.

Restrict the null-pointer check to n > 0, matching the convention already
used by Parallel_Grid::reduce (only a null buffer with a non-zero size is
a genuine bug). n < 0 remains a hard error. Regression introduced in
d9685d4eb when the inline packing loops were extracted into these helpers.

* Refactor: move rhog_io into module_charge as chg_rhog_io

Relocate source_estate/rhog_io.{h,cpp} to source_estate/module_charge/
under the module_charge namespace, rename include guard to CHG_RHOG_IO_H,
and update the warning tags emitted at runtime. Update both callers
(chg_init.cpp, esolver_fp.cpp) and build files; adapt test_rhog_io.cpp in
place ahead of its move in a follow-up commit. No behavior change.

* Refactor: create module_charge/test with the rhog io unit test

Move test_rhog_io.cpp into module_charge/test/test_chg_rhog_io.cpp with
its support data charge-density.dat, register the new test subdirectory,
and rename the target to MODULE_CHARGE_rhog_io. Remove the migrated
AddTest block from the legacy source_estate/test/CMakeLists.txt.

* Refactor: move charge and charge-extra unit tests into module_charge/test

Rename charge_test.cpp to test_charge.cpp and charge_extra_test.cpp to
test_chg_extra.cpp per the test naming rule, move prepare_unitcell.h
alongside its only users, and register MODULE_CHARGE_charge /
MODULE_CHARGE_extra in the module_charge test CMakeLists. No test data
moves: prepare_unitcell.h only sets file-name strings at runtime, and
the extra test only writes cube files into ./support/.

* Refactor: move mix, parallel and tools unit tests into module_charge/test

Relocate test_chg_mix.cpp (fixing its relative includes), test_chg_parallel.cpp
and test_chg_tools.cpp into module_charge/test, register MODULE_CHARGE_tools /
MODULE_CHARGE_mix / MODULE_CHARGE_parallel with the 4-process mpirun test,
and drop the migrated blocks from the legacy source_estate/test CMakeLists.

* Refactor: rename module_charge test dir to unittests and wire CI for it

Rename source_estate/module_charge/test to unittests (relative CMake
paths are immune to the move). Sync the referencing points: the
add_subdirectory call, the coverage lcov filter (add '*/unittests/*' so
test sources stay excluded from the report), a dedicated Module_Charge
ctest step in test.yml with MODULE_CHARGE added to the catch-all -E
list to avoid double execution, and unittests/ added to the
code_quality_score.py SKIP_DIRS.

* Fix: pass ucell.omega to Charge::sum_rho/renormalize_rho to fix NPT stress

Root cause: commit 34b441e1c ("Refactor: remove redundant Charge::omega_
pointer") changed Charge::sum_rho() to read the cell volume from
rhopw->omega instead of ucell.omega. In variable-cell calculations (NPT),
pw_rho/pw_rhod are NOT rebuilt on cell change (only pw_wfc is), so
rhopw->omega keeps the initial cell volume while ucell.omega is updated
every MD step. The stale volume made sum_rho() return a wrong electron
count, which made renormalize_rho() scale rho by the wrong factor,
corrupting the stress (deviation ~0.002 in 095_PW_NPT) while the total
energy stayed near-correct (variational, second-order sensitive).

Fix: add an explicit omega parameter to Charge::sum_rho() and
renormalize_rho(); all call sites (init_scf, chg_routine, LCAO dm2rho path
through HSolverLCAO/dmToRho, RDMFT update_charge, OFDFT renormalize_psi)
now pass ucell.omega. This mirrors the existing check_rho(..., ucell.omega)
pattern.

Also mark three other rhopw->omega users with BUG(investigate) comments:
get_local_pp_energy, cal_delta_escf, and Makov-Payne correction. These are
pre-existing and were not changed by the refactor; they may have the same
stale-volume issue in NPT and should be investigated separately.

Bisected to 34b441e1c over the 20260916 module_charge refactor branch.

* Fix: add omega arg to remaining dm2rho call sites

Missed four LCAO_domain::dm2rho call sites in the previous commit:
- lcao_set.cpp init_chg_dm (skip_normalize=true, omega unused)
- esolver_dm2rho.cpp
- esolver_ks_lcao_tddft.cpp weight_dm_rho
- module_dm/init_dm.cpp

All now pass ucell.omega.

* Fix: restore HamiltHSMatrix hs declaration in cal_mw_from_lambda

Accidentally removed the line while editing the comment.

* Fix: close_kerker_gg0 actually disables Kerker; drop dead mixing_gg0 members

The chg_precond refactor (commit 6d127d517) made the Kerker kernels read
cfg_ (immutable INPUT snapshot) instead of Charge_Mixing members, but
close_kerker_gg0() kept writing the now-dead mixing_gg0/mixing_gg0_mag
members. As a result, the non-separate-loop EXX path in exx_lri_interface.hpp
silently failed to disable Kerker after convergence.

Fix: add a kerker_disabled_ flag on Charge_Mixing that the mix_rho_recip/
mix_rho_real screening lambdas short-circuit on. The flag lives on the
object, not in cfg_, so the immutable INPUT snapshot invariant is preserved.

Also drop the now-dead members mixing_gg0/mixing_gg0_mag/mixing_gg0_min/
mixing_angle/mixing_dmr and the get_mixing_gg0() getter; set_mixing/init_mixing
now read these from cfg_ directly. Add CloseKerkerGg0DisablesScreenReal
regression test that compares close_kerker_gg0() output against the
cfg.mixing_gg0=0 baseline and proves the flag is load-bearing.

* Fix: relax over-strict null-buffer asserts for empty grid partitions

reduce_diff_pools and Parallel_Grid::reduce_across_pools still forbade
null buffers unconditionally, contradicting the rule documented at
parallel_grid.cpp:355-360. A rank with nrxx == 0 may legitimately hold
a null rho/kin_r pointer; the MPI calls below use count 0 and ignore
the buffer. Align both call sites with the documented rule.

* Fix: relax over-strict null-buffer assert in ParaRgridWorld::reduce_across_pools

Same pattern as the previous fix: a rank with nrxx == 0 legitimately
holds a null buffer, and MPI_Allreduce with count 0 ignores it.
Align with the rule documented at parallel_grid.cpp:355-360.

* Fix: allow nnr == 0 in DMR mixing for empty MPI partitions

nnr is local to each MPI rank and may legitimately be zero when no
atom pairs survive the cutoff on that rank. The previous check
aborted DMR mixing for such distributions, whereas the historical
implementation allowed empty blocks. Relax the guard in
check_dmr_inputs() and init_mixing_dmr() to reject only negative
nnr, and require non-null DMR buffers only when nnr > 0, matching
the established nrxx == 0 convention in module_charge.

* Fix: split reciprocal rho copy from real-space |m| rescale in mix_rho_recip

The nspin==4 && mixing_angle>0 branch of mix_rho_recip mixed two
distinct operations in one loop bounded by npw, but rho_magabs is
sized nrxx (real-space) and the new |m| is written back by
recip2real into rho_magabs[0..nrxx-1]. Reading rho_magabs[npw+ig]
goes out of bounds once npw+ig >= nrxx (AddressSanitizer reproduces
with nrxx=125, npw=93) and the loop bound npw leaves the real-space
tail [npw, nrxx) of {mx,my,mz} unscaled. Split into two loops: the
reciprocal rho copy stays bounded by npw, the magnetization rescale
is bounded by nrxx and reads rho_magabs[ir].

* Refactor: remove unused Charge_Mixing::conserve_setting

conserve_setting() was introduced by 420f1ad00 (DeltaSpin feature
merge, 2026-06-15) but never wired up: no production caller, no
test reference, and the DeltaSpin module does not touch
Charge_Mixing. Drop the dead declaration per the project rule that
unused functions and their tests be removed.

* Refactor: drop dead Charge_Mixing::tpiba2 member

tpiba2 was declared in chg_mix.h but never assigned by set_mixing()
nor read anywhere in the module. Grep across the whole source tree
confirms all tpiba2 references are either ucell.tpiba2 (a separate
UnitCell member) or local variables in unrelated modules. The
Charge_Mixing class never computed or used its own tpiba2 pointer;
only tpiba is consumed by the stateless Kerker kernels via
mix_rho_recip/mix_rho_real. Remove the dead declaration.

* Refactor: route Charge_Mixing getters through cfg_

get_mixing_mode(), get_mixing_beta(), get_mixing_ndim() previously
returned the legacy mirror members that set_mixing() kept in sync
with cfg_ by hand. With cfg_ now treated as the immutable INPUT
snapshot, route the public getters through cfg_ directly so there
is a single source of truth for INPUT parameters. External callers
(esolver_ks_lcao, lcao_others, pw_others) are unaffected since
signatures are unchanged. The legacy members remain in place for
now; they are dropped in a later step after internal readers are
migrated.

* Refactor: init_mixing constructs Mixing from cfg_ not legacy mirrors

init_mixing() branched on this->mixing_mode and passed
this->mixing_ndim/mixing_beta to the Broyden/Pulay/Plain_Mixing
constructors. These legacy mirrors were kept in sync with cfg_
manually by set_mixing(). Route through cfg_ directly so cfg_
remains the single source of INPUT parameters. The Mixing objects
themselves still copy beta/ndim into their own members at
construction; that is a one-time snapshot and not a continuous
sync surface, so it is left untouched.

* Refactor: mix_rho_recip/mix_rho_real read mixing_beta from cfg_

Both mix_rho_recip and mix_rho_real built the twobeta_mix functor by
reading this->mixing_beta / this->mixing_beta_mag, which are legacy
mirrors that set_mixing() kept in sync with cfg_. Route the six
construction sites through cfg_.mixing_beta / cfg_.mixing_beta_mag
so cfg_ is the single source of INPUT parameters consumed by the
mixing logic. Behavior is unchanged since the mirrors and cfg_
hold identical values after set_mixing().

* Refactor: set_mixing stops mirroring cfg_ into legacy members

set_mixing() copied mixing_mode, mixing_beta, mixing_beta_mag,
mixing_ndim from cfg into legacy mirror members, then validation
and logging read from the mirrors. Now that all internal readers
(init_mixing, mix_rho_recip, mix_rho_real, getters) read from
cfg_, the mirror writes are dead work. Drop them and route
validation and log output through cfg_ directly. omega and tpiba
remain pointer members because they alias external runtime state
(cell volume, lattice constant) that changes across SCF iterations
and so do not belong in MixingConfig (an immutable INPUT snapshot).

* Refactor: drop legacy Charge_Mixing mirror members; cfg_ is single source

Drop mixing_mode, mixing_beta, mixing_beta_mag, mixing_ndim mirror
members. After the previous commits every internal reader (getters,
init_mixing, mix_rho_recip, mix_rho_real, set_mixing validation
and log output) routes through cfg_, so the mirrors are dead state
that set_mixing() no longer writes. cfg_ is now the single source
of truth for INPUT mixing parameters.

Update test_chg_mix.cpp accordingly: the two assertions that
reached directly into CMtest.mixing_beta_mag and CMtest.mixing_mode
now read CMtest.get_mixing_config().mixing_beta_mag and
CMtest.get_mixing_mode(), matching the public API used by the
other assertions in the same block. No production caller accessed
these members directly (esolver_ks_lcao, lcao_others, pw_others
all used the getters), so the change is test-only on the consumer
side.

* Refactor: drop NSDMI from MixingConfig to force explicit construction

The non-static data member initializers in MixingConfig provided
plausible-looking defaults (e.g. mixing_beta=0.8, mixing_mode=
"broyden") that silently masked forgotten fields when a new field
was added but not wired up at construction sites. With the
defaults removed, every construction site must use aggregate
initialization (or copy-assign from a fully-initialized instance),
and a missing field yields value-initialized (zero/empty) members
that are far more likely to trip a test than the old defaults.
Combined with -Wmissing-field-initializers promoted to error in
the next commits, adding a field to MixingConfig without updating
all aggregate-initialization sites becomes a compile error.

* Refactor: aggregate-init MixingConfig in esolver_ks with pragma guard

Convert the 17-line field-by-field assignment of mix_cfg into a
single aggregate initialization in declaration order. Wrap it in
#pragma GCC diagnostic error "-Wmissing-field-initializers" so
that adding a field to MixingConfig without updating this list
becomes a compile error rather than silently using a default.
Each initializer is annotated with the field name it corresponds
to, making the declaration-order dependency auditable at a glance.

* Refactor: aggregate-init MixingConfig in test_chg_mix with pragma guard

Convert make_cfg()'s 17-line field-by-field assignment into a
single aggregate initialization in declaration order, matching
the esolver-side change. Wrap in the same
#pragma GCC diagnostic error "-Wmissing-field-initializers" so
that adding a field to MixingConfig without updating the test
helper is also a compile error. Both construction sites (esolver
and test) now fail at compile time if a field is missing, closing
the maintenance gap where a new field could silently fall back to
a default value.

* Fix: fail-fast guards in Charge_Mixing and update chg_mix tests

Add validation to turn latent misuse (skipped set_rhopw/set_mixing)
into clear WARNING_QUIT errors instead of null dereference or heap
corruption:
- init_mixing rejects a null rhopw
- if_scf_oscillate checks scf_nmax > 0 and iteration range
- mix_rho validates chr/chr->rhopw and the grid pointers

Fix three chg_mix unit tests that read cfg_ before set_mixing, which
caused a SIGSEGV in SCFOscillationTest and assertion failures in the
two inner-product tests.

* test(module_charge): add unit tests for chg_uspp and chg_dmr

Add test_chg_uspp.cpp covering split_dgrid/merge_dgrid (normal split,
round-trip, nspin=1/2, empty high-frequency/smooth boundaries, and
input-validation abort paths).

Add test_chg_dmr.cpp covering init_mixing_dmr/mix_dmr (nspin=1/2/4
mixing with Plain_Mixing analytically verified, empty-partition null
buffer allowance, and input-validation abort paths).

Wire both targets into unittests/CMakeLists.txt.

* test(module_charge): add unit tests for chg_precond, chg_drho, chg_drho_inner, chg_mix_rho

- test_chg_precond.cpp: kerker_screen_recip/real (early return, nspin=1/2/4
  filter, nspin=4 with mixing_angle resize, real-space matches reciprocal).
- test_chg_drho.cpp: inner_product_real, cal_drho real-space path
  (nspin=1/2/4+domag_z), cal_dkin (meta_gga false/true).
- test_chg_drho_inner.cpp: inner_product_recip_rho and
  inner_product_recip_hartree for nspin=1 with a single G component,
  analytically verified against the Coulomb metric.
- test_chg_mix_rho.cpp: mix_rho abort paths (null chr/chr->rhopw, unset
  rhopw, double_grid without rhodpw) and real-space plain mixing value.

Wire all four targets into unittests/CMakeLists.txt.

* test(module_charge): add unit tests for chg_symm, chg_symm_detail, chg_atomic, chg_atomic_inner

- test_chg_symm.cpp: symmetrize_rho / cal_rhog_symm / cal_rhog_symm_soc
  no-op paths when symm_flag == 0, for nspin=1 and nspin=4.
- test_chg_symm_detail.cpp: psymmg and psymmg_soc idempotence on a
  manually built D_4 point group over a serial cubic PW_Basis.
- test_chg_atomic_inner.cpp: compute_rhoatm USPP direct-copy branch and
  NCPP integrate+scale-to-zv branch (Gaussian rho_at with known analytic
  integral); normalize_and_check renormalizes uniform density to nelec.
- test_chg_atomic.cpp: atomic_rho ntype==0 path (skips atom loop) and
  spin_number_need==3 abort path.

Wire all four targets into unittests/CMakeLists.txt.

* test(module_charge): add chg_tau/chg_routine/chg_init tests; drop spurious XC_Functional stubs

Fourth batch of module_charge unit tests:
- test_chg_tau.cpp: mix_tau_recip abort paths (null chr/grid/mixing, nspin<1,
  double_grid without high-f mixer) and non-double-grid plain mixing value.
- test_chg_routine.cpp: chgmixing_ks_pw/lcao iter==1 restart-step setup, and
  chgmixing_ks convergence branches (conv_esolver true / drho<hsolver_error
  skip mix_rho).
- test_chg_init.cpp: init_rho "wfc" with null wfcpw abort, and "atomic" with
  ntype==0 + meta_gga Thomas-Fermi tau initialization.
Wire all three targets into unittests/CMakeLists.txt.

Cleanup: remove the XC_Functional::func_type / ked_flag definitions from
test_chg_drho, test_chg_mix_rho, test_chg_symm, test_chg_atomic_inner,
test_chg_atomic, test_chg_tau, test_chg_routine, and test_chg_init. None of
the non-test sources compiled into these targets reference these statics
(charge.cpp, chg_*.cpp, and the linked base/cell_info/planewave_serial/
symmetry libraries are clean), so the definitions were pure dead weight.
Also drop the now-unneeded xc_functional.h include from test_chg_mix_rho.cpp
and correct the stub comments.

* test(module_charge): fix broken includes in unit tests

- include chg_atomic_detail.h instead of nonexistent chg_atomic_inner.h
  in test_chg_atomic_inner.cpp; add math_integral.h for Simpson_Integral
- include chg_drho.h in test_chg_drho_inner.cpp for
  module_charge::inner_product_recip_hartree
- include source_cell/magnetism.h in tests that define Magnetism stubs
  (test_chg_drho, test_chg_symm, test_chg_tau, test_chg_mix_rho)
- fix nonexistent source_charge/mixing includes in test_chg_tau.cpp to
  source_base/module_mixing

* test(module_charge): fix link/build issues; temporarily disable routine/init targets

- test_chg_tau: use Plain_Mixing(beta) ctor and init_mixing_data with
  complex type_size (old set_mixing_beta/init_mixing no longer exist)
- test_chg_symm_detail: add Magnetism stub required by cell_info's
  unitcell.cpp, matching other tests in this directory
- test_chg_routine: adapt to two-arg set_rhopw and tpiba from ucell
- disable MODULE_CHARGE_routine and MODULE_CHARGE_init targets with
  documented reasons: their transitive dependencies (Plus_U_Base,
  elecstate, source_io) are deeply coupled; to be resolved later

* fix bug

* format tool_quit

* remove a test due to WARNING_QUIT funcitno

* delete the support file charge-density.dat because unittests never need it

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
Co-authored-by: Xiaoyang Zhang <tsfxwbbzxy@163.com>
Redo of the work in #7988 and #7990, both of which were closed while the
charge density module was being restructured. That restructuring landed in
#7972 and already did most of the decoupling those PRs proposed: allocate(),
renormalize_rho() and sum_rho() now take their inputs explicitly, the mixing
parameters are aggregated in a MixingConfig, and chg_mix.cpp / chg_drho.cpp /
charge.cpp are free of global parameter reads. What was left was the
test-side access.

Production changes are additive only - no existing signature moves and no
line is deleted from any production header:

  Charge::get_allocate_rho()          - report whether allocate() has run
  Charge_Mixing::get_rho_mdata()      - mirror the existing get_dmr_mdata()
  Charge_Mixing::get_tau_mdata()
  Charge_Mixing::set_mixing_config()  - pair for the existing getter, for
                                        callers that must update the snapshot
                                        without rebuilding the mixing history
  XC_Functional::set_func_type()      - pair for get_func_type()
  XC_Functional::set_ked_flag()       - pair for get_ked_flag()

Test changes:

  test_dm_r_init   - two sites move to the already public get_DMR_save()
  test_charge      - the global parameter scratchpad becomes fixture state
                     (32 refs -> 0); PW_Basis setup goes through the public
                     initgrids/initparameters/setuptransform sequence instead
                     of the protected distribute_r()/distribute_g()
  test_chg_mix     - the scratchpad becomes a fixture-owned MixingConfig
                     (163 refs -> 0); the three blocks that hand-wired
                     Charge::_space_* now take their buffers from the fixture,
                     which owns them as vectors and points the public
                     rho/rhog/kin_r views at them with the same stride

No expected value or tolerance was changed.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* Fix: Correct FFTW version detection and NCCL build summary

* Read FFTW version directly from adjacent pkg-config metadata
* Remove unneeded README20260902 file

* Fix build without LibRI: split BvK utils out of ri_util.h

module_lr is built whenever ENABLE_LCAO is on, but lr_io_krlist.cpp
unconditionally included module_ri/ri_util.h, which pulls in LibRI
headers (<RI/global/Array_Operator.h>) and fails to compile when
ENABLE_LIBRI is off (regression from #7849).

Move the LibRI-free Born-von Karmen helpers (get_Born_vonKarmen_period,
get_Born_von_Karmen_cells) into a new header ri_util_bvk.h; ri_util.h
now includes it, and lr_io_krlist.cpp includes only the new header.

Verified: target lr builds with ENABLE_LIBRI=OFF (build/), target ri
builds with ENABLE_LIBRI=ON (build_std_para/).

* Fix timer_enable_nvtx: define __USE_NVTX on the targets that consume it

__USE_NVTX was defined only on the final executable target, whose sole
translation unit main.cpp contains no NVTX code. The two OBJECT libraries
that actually guard NVTX calls with the macro -- base (source_base/timer.cpp)
and driver (source_main/driver.cpp) -- never saw it, so every NVTX block was
preprocessed away and timer_enable_nvtx had no effect in any CUDA build.

Move the definition onto base and driver, and link CUDA::nvToolsExt for
CUDA toolkits older than 12.9 (NVTX is header-only since 12.9).

Verified with build_pw_gpu (USE_CUDA=ON): base/driver targets compile with
NVTX symbols present in timer.cpp.o, driver.cpp.o references
timer::enable_nvtx_, and the full abacus_pw_gpu executable links
(v3.11.0-beta9).

* Refactor charge density module (#7972)

* module_charge: normalize indentation and brace single-statement control flow

Mechanical cleanup as the first step of the module_charge governance
refactor: convert leading tabs to 4-space indentation (1011 occurrences
across 11 files) and add braces around all single-statement if/for/while
bodies (11 sites). No functional change.

* module_charge: aggregate Charge_Mixing params into MixingConfig

Introduce a MixingConfig POD that bundles the INPUT mixing parameters
with the runtime globals (nspin, scf_thr_type, double_grid), and change
set_mixing from a 12-argument interface to set_mixing(const MixingConfig&,
double&, double&). Charge_Mixing now stores the config and reads nspin /
scf_thr_type / double_grid from it instead of PARAM.inp / PARAM.globalv,
removing the direct PARAM reads in set_mixing and init_mixing.

The single production call site (esolver_ks.cpp) fills the config, and
the unit test drives set_mixing via a make_cfg() helper. The
'#define private public' access hack is kept for now with a TODO: the
test still must write Parameter::input/sys, Charge::_space_* and
XC_Functional privates, which need the Step 4/5 global-state
parameterization before it can be removed.

Verified: make -j30 MODULE_ESTATE_charge_mixing (build_max_para_test)
passes with no errors.

* module_charge: deduplicate twobeta_mix lambdas and replace raw new with std::vector

Extract the repeated two-beta mixing functor in mix_rho_recip/mix_rho_real
into a make_twobeta_mix<T> template helper (6 lambda copies removed), and
convert all local raw new[]/delete[] buffers in charge_mixing_rho.cpp to
zero-initialized std::vector, dropping the paired ZEROS calls.

* module_charge: move residual/inner-product globals into MixingConfig

Extend MixingConfig with gamma_only_pw/domag/domag_z so mix_resid.cpp
(get_drho, get_dkin, inner_product_recip_{rho,simple,hartree,real}) no
longer reads PARAM/GlobalV; all branches now consume this->cfg_.
inner_product_recip_rho's raw pointer-array views are switched to
std::vector. Production fills the three new fields in esolver_ks, and
the test fixture gains a sync_cfg() helper to push PARAM mutations into
cfg_ for the inner-product branch tests.

* module_charge: own Charge's _space_* storage with std::vector (Step 5a)

Replace the six private raw _space_rho/_space_rho_save/_space_rhog/
_space_rhog_save/_space_kin_r/_space_kin_r_save buffers with
std::vector, so Charge's underlying contiguous storage self-manages and
the matching delete[] calls in destroy() (which relied on reading
possibly-uninitialized pointers) go away. The public rho/rhog/rho_save/
rhog_save/kin_r/kin_r_save views keep their double**/complex** shape and
still alias the vector memory via .data(), so all external consumers are
unaffected. Tests that drove _space_* directly are adapted to
resize()/.data() and drop their manual delete[] of the buffers.

* module_charge: route chgmixing_ks through its inp parameter

chgmixing_ks already takes a const Input_para& inp but still read
PARAM.inp.mixing_restart / PARAM.inp.scf_nmax from the global. Use the
inp argument instead so the function no longer reads INPUT state through
the global for these two fields. PARAM.globalv.ks_run is a runtime
per-process flag (set from band-parallel topology), not an input, so it
is intentionally left as-is rather than threading it through the
interface.

* module_charge: split Charge::init_rho into per-stage private methods

init_rho had a cyclomatic complexity of 36 from five sequential stages
(file read, atomic fallback, Thomas-Fermi tau, restart load, wfc read)
interleaved through shared read_error/read_kin_error flags. Extract the
four branches into private methods -- read_rho_from_file,
init_rho_atomic_and_tau, load_rho_from_restart, init_rho_from_wfc -- and
leave init_rho as a thin sequence of stage calls. Logic is unchanged; the
error flags are threaded through as parameters. The deepest stage
(read_rho_from_file) now sits at complexity 19, down from 36 for the
monolith. The remaining global reads inside the stages are untouched and
deferred to a later parameterization step.

* module_charge: extract Charge density math into charge_math free functions

sum_rho, cal_rho2ne and non_linear_core_correction each used Charge
members only to reach a handful of scalars (nrxx/nxyz/omega) or the
reciprocal-shell table (gg_uniq/ngg); the rest of each body is pure
numerics. Move the three bodies into a new charge_math namespace as free
functions with those values passed explicitly, and leave the Charge
members as thin forwarding wrappers so no caller outside the module
changes. The kernels are now unit-testable in isolation and no longer
coupled to Charge state. One behavior note: the pre-quit debug line that
printed sum_rho to ofs_warning is dropped so the free function stays free
of global-stream dependencies. charge_math.cpp is wired into the estate
library and the charge_test target.

* module_charge: register charge_math.o in the hand-written Makefile build

The CMake build already picks up charge_math.cpp; mirror that in
Makefile.Objects so the legacy Makefile flow links the new charge_math
kernels too. The module_charge directory is already on VPATH, so adding
charge_math.o to the object list is sufficient.

* module_charge: extract Charge::atomic_rho into charge_atomic free function

Remove Charge::atomic_rho entirely and replace all call sites with
module_charge::atomic_rho(..., rhopw), eliminating the need for a thin
wrapper on the Charge class. This decouples atomic density initialization
from Charge's state and improves charge.cpp quality score from 2 to 44.

* module_charge: forbid Charge copies and guard tau.cube write

scf_out_chg_tau aborted in Parallel_Grid::reduce on
assert(rhoin != nullptr) because the kin_r_save[is] handed to
write_vdata_palgrid was not a valid buffer. After the _space_* storage
became std::vector (ecf5084d4), a copied/moved Charge leaves its
rho/kin_r views dangling into another object's vector buffer, and a
kin_r_save never allocated (ked_flag set after allocate) stays nullptr;
both surface as a null rhoin deep inside MPI gather instead of at the
source.

Delete Charge's copy constructor/assignment so any value copy of the
vector-aliasing views fails at compile time, and check kin_r_save in
ctrl_output_fp before writing tau.cube so a missing allocation reports
a clear message instead of tripping the MPI assert.

Verification: not run locally (per user request, user compiles).

* module_base: tolerate null grid buffer when a rank owns no grid points

scf_out_chg_tau (LCAO, SCAN, out_chg=1, 4 MPI ranks) aborted in
Parallel_Grid::reduce on assert(rhoin != nullptr). Bisecting between
83eb5d0f3 (good) and ecf5084d4 (bad) isolated the regression to
ecf5084d4, which moved Charge's _space_* storage from raw new[] to
std::vector.

Root cause: with 4 ranks the FFT grid is slab-decomposed so that the
last rank owns zero real-space points (nrxx == 0, confirmed via a
temporary diagnostic printing fn/is/rank/nrxx at the reduce call site).
Before ecf5084d4, _space_rho = new double[nspin * 0] == new double[0]
returned a unique non-null pointer, so rho_save[is] was non-null and the
assert passed. After the change, an empty vector's .data() returns
nullptr, so the rank with nrxx == 0 handed a null rhoin to reduce and
tripped the assert (Debug) or fed MPI_Gatherv a null buffer (Release).

A rank with nrxx == 0 is legitimate: MPI_Gatherv is invoked with
sendcount 0 and ignores the send buffer. Relax the assert to only flag a
null buffer when nrxx != 0, and revert the now-unneeded kin_r_save guard
in ctrl_output_fp (it would have falsely aborted on the nrxx == 0 rank).

Verification: Release build (build_max_para_test), ran
  cd tests/03_NAO_multik/scf_out_chg_tau &&
  OMP_NUM_THREADS=1 mpirun -np 4 ../../../build_max_para_test/abacus_max_para
Result: exit 0, chg.cube and tau.cube written; numerical comparison
against chg.cube.ref/tau.cube.ref gives maxdiff 0 (chg) and 1e-14 (tau).

* module_charge: extract Charge::set_rho_core into charge_math free function

Move set_rho_core to charge_math::set_rho_core with rho_core,
rhog_core and rhopw passed explicitly instead of reading Charge
state, and call charge_math::non_linear_core_correction directly.
Remove the now-unused Charge::non_linear_core_correction wrapper,
use std::vector for the rhocg/vg scratch buffers, update the
init_scf call site, and drop the obsolete member stubs in the
elecstate unit tests.

* module_charge: vectorize Charge_Extra history arrays and forbid copies

Replace the raw new[]/delete[] displacement arrays (dis_old1, dis_old2,
dis_now) with std::vector and remove the hand-written destructor. This
fixes a read of uninitialized pot_order when an object is destroyed
before Init_CE, a memory leak when Init_CE is called repeatedly, and a
double-free risk from the implicitly generated shallow copy. The copy
constructor and copy assignment are deleted so the molecular-dynamics
trajectory history cannot be silently forked. The unit test now checks
vector sizes instead of non-null pointers.

* Rename charge_math to chg_tools and unify namespace module_charge

- Rename module_charge/charge_math.{h,cpp} to chg_tools.{h,cpp} via git mv
- Change namespace charge_math to module_charge to match charge_atomic
  and chgmixing in the same directory
- Update include guard CHG_TOOLS_H and TITLE/timer labels accordingly
- Update call sites in init_scf.cpp, charge.cpp, charge_init.cpp
- Update build references in Makefile.Objects and both CMakeLists.txt

* module_charge: refactor Symmetry_rho class to free functions

Convert the stateless class Symmetry_rho into namespace module_charge
free functions and rename files for consistency:
  symm_rho.{h,cpp}      -> chg_symm.{h,cpp}
  symm_rho_detail.h     -> chg_symm_detail.h
  symm_rhog.cpp         -> chg_symm_detail.cpp

- 5 public functions become module_charge::symmetrize_rho / cal_rhog_symm
  (2 overloads) / cal_rhog_symm_soc (2 overloads)
- 2 cross-TU helpers (psymmg/psymmg_soc) moved to module_charge::detail
  via chg_symm_detail.h
- 3 internal MPI helpers moved to anonymous namespace
- Delete dead code psymm (real-space symmetrization, never called)
- Remove empty ctor/dtor and parallel_grid.h include
- Rename begin/begin_soc to cal_rhog_symm/cal_rhog_symm_soc for clarity
- Update timer/TITLE labels from "Symmetry_rho" to "module_charge"
- Migrate all 14 call sites and 1 test stub
- Remove obsolete Makefile special rule (no more name collision)

* module_charge: extract MixingConfig header and drop unused inner_product_recip_simple

Move MixingConfig from charge_mixing.h into its own mixing_config.h so
stateless residual kernels can include the config without dragging in
Charge_Mixing. Remove inner_product_recip_simple, which had no production
call sites, together with its unit test.

* module_gint: move gint_prec_ctrl from module_charge

Relocate gint_prec_ctrl.{h,cpp} and its test into module_gint, update the
include in esolver_ks_lcao.h and rewire the CMake/Makefile object lists.

* module_charge: extract mixing inner products into chg_drho free functions

Rename mix_resid.cpp to chg_drho.cpp and turn inner_product_real and
inner_product_recip_hartree into module_charge free functions declared
in chg_drho.h; inner_product_recip_rho, which is only shared with the
unit test, moves to module_charge::detail in chg_drho_detail.h.
Charge_Mixing loses the three private inner-product members and
mix_rho_recip/mix_rho_real bind the free functions through lambdas.
get_drho/get_dkin stay as members for this step.

* module_charge: hide cal_drho/cal_dkin in an anonymous namespace

Move the get_drho/get_dkin implementations into file-local cal_drho/
cal_dkin free functions with all inputs explicit; the public
Charge_Mixing methods become thin forwarding wrappers so esolver call
sites stay unchanged.

* module_gint: fix include path in test_gint_prec_ctrl after relocation

* module_charge: extract Kerker screen kernels into chg_precond free functions

Move Charge_Mixing::Kerker_screen_recip/real to module_charge namespace
as free functions in chg_precond.{h,cpp}, renaming mix_precond.cpp via
git mv. Config/grid/geometry are passed explicitly via MixingConfig,
PW_Basis*, and tpiba, eliminating the function's direct read of
PARAM.inp.nspin. Replace 8 std::bind call sites in charge_mixing_rho.cpp
with lambdas, update 2 commented-out bind sites in charge_mixing_dmr.cpp,
and rewrite 12 test call sites in charge_mixing_test.cpp to construct an
independent MixingConfig instead of poking at Charge_Mixing privates.
Drop the now-unused member function declarations from charge_mixing.h.

* module_charge: fix Makefile.Objects after mix_precond -> chg_precond rename

Update the non-CMake object list to track the renamed translation unit so
make-based builds do not reference the deleted mix_precond.o.

* module_charge: drop Charge_Mixing::get_drho/get_dkin wrappers

Expose cal_drho/cal_dkin as module_charge free functions in chg_drho.h
and let ESolver_KS call them directly with explicit arguments; add
Charge_Mixing::get_mixing_config() as a const observer for the config.

* module_charge: rename chgmixing.h/cpp to chg_routine.h/cpp

Align with the chg_<feature> naming pattern used in the same directory
(chg_drho, chg_precond, chg_symm, chg_tools). Update include guard to
CHG_ROUTINE_H, the self-include in chg_routine.cpp, the entry in
source_estate/CMakeLists.txt and source/Makefile.Objects, and the three
#include sites in esolver_ks{,_pw,_lcao}.cpp. Function names
(chgmixing_ks{,_pw,_lcao}) and TITLE/timer tags are intentionally left
unchanged to keep the diff minimal.

* module_charge: rename mixing_config.h to chg_mix_cfg.h

Rename the MixingConfig header to align with the chg_* naming
convention in module_charge. Update the include guard and the four
in-tree includers; no CMake change is needed since the header is not
listed explicitly.

* module_charge: convert Charge MPI helpers into chg_parallel free functions

Rename charge_mpi.cpp to chg_parallel.cpp and add chg_parallel.h, moving
the three stateless Charge member functions (reduce_diff_pools, rho_mpi,
kin_r_mpi) to module_charge namespace free functions that take the
Charge object explicitly. Remove their declarations from charge.h and
update all call sites in elecstate_pw, stress_mgga, read_wf2rho_pw and
sto_iter. Rename the unit test to test_chg_parallel.cpp and update the
test target name accordingly.

GlobalV/PARAM reads and the direct MPI_Allreduce in reduce_diff_pools
are preserved as pre-existing technical debt (migration-neutral).

* Rename charge_atomic files to chg_atomic

- Rename module_charge/charge_atomic.{h,cpp} to chg_atomic.{h,cpp}
- Update include guard to CHG_ATOMIC_H
- Update includes in charge_init.cpp and charge_extra.cpp
- Update source paths in CMakeLists.txt, test CMakeLists.txt
- Fix stale object names in Makefile.Objects: replace
  symm_rho_charge.o/symm_rhog.o with chg_symm.o/chg_symm_detail.o

* module_charge: extract USPP double-grid split/merge into chg_uspp free functions

Introduce module_charge::split_dgrid / merge_dgrid in chg_uspp.{h,cpp} as
RAII, parameter-explicit replacements for Charge_Mixing::divide_data /
combine_data / clean_data, which paired raw new[] with manual delete[]
across ~160 lines of mixing code.

- chg_uspp.{h,cpp}: stateless free functions in module_charge namespace;
  outputs are caller-pre-sized std::vector, no new/delete; parameter
  validation via WARNING_QUIT; TITLE/timer tags preserved
- charge_mixing_rho.cpp: rho and tau double-grid paths switched to the new
  functions; raw pointer aliases kept for !double_grid so the existing
  mixing call sites (nspin==1/2/4) are untouched
- CMakeLists.txt (source + test): wire chg_uspp.cpp

The legacy divide_data/combine_data/clean_data members are not yet removed;
that follows in a later step after the test is updated.

* module_charge: rewrite MixDivCombTest for the new split_dgrid/merge_dgrid

Drop the legacy alias-pointer assertions (EXPECT_EQ(datas, data.data()),
EXPECT_EQ(datas, nullptr) after clean_data) that coupled the test to the
old new[]/delete[] ownership model.

The rewritten case verifies the actual contract:
- split_dgrid fills smooth and high-frequency buffers with the dense
  data verbatim (per-element comparison)
- merge_dgrid is a left-inverse of split_dgrid (output == input)
- no explicit cleanup call is required: std::vector manages storage

Covers nspin == 1 and nspin == 2 paths.

* module_charge: drop legacy divide_data/combine_data/clean_data members

With the new module_charge::split_dgrid/merge_dgrid in chg_uspp.{h,cpp}
and all call sites in charge_mixing_rho.cpp migrated, the original
Charge_Mixing::divide_data / combine_data / clean_data members are dead.

- delete charge_mixing_uspp.cpp (the raw new[]/delete[] implementation)
- drop the three member declarations from charge_mixing.h
- remove charge_mixing_uspp.cpp from source/test CMakeLists.txt
- Makefile.Objects: drop charge_mixing_uspp.o, add chg_uspp.o
- refresh one stale comment in charge_mixing_rho.cpp to reference
  merge_dgrid instead of the removed combine_data

* module_charge: rename charge_extra files to chg_extra and move class into namespace

Rename charge_extra.h/cpp to chg_extra.h/cpp and wrap the Charge_Extra
class in the module_charge namespace, matching the rest of module_charge
(chg_atomic, chg_symm, chg_uspp). Update include guards, call sites in
esolver_fp.h and the unit test, and CMake/Makefile source lists.

* module_charge: extract DMR mixing into chg_dmr free functions

Move the DMR allocation/mixing logic out of Charge_Mixing members into
stateless module_charge functions (init_mixing_dmr, template mix_dmr
with explicit instantiation), passing the Mixing object, mixing data
and MixingConfig explicitly instead of reading PARAM. Merge the two
identical real/complex mix_dmr overloads, replace raw new[]/delete[]
of the magnetic buffers with std::vector, and de-duplicate the
two-beta mixing lambda into a file-local helper. The members stay as
thin timer-wrapped wrappers so external call sites are unchanged.

* module_charge: remove Charge_Mixing DMR wrappers, call chg_dmr directly

Delete charge_mixing_dmr.cpp and have the two call sites
(chg_routine.cpp, esolver_ks_lcao.cpp) invoke module_charge::
init_mixing_dmr/mix_dmr directly with the Mixing object, mixing data
and MixingConfig obtained through Charge_Mixing accessors. Expose the
owned DMR mixing history via a new get_dmr_mdata() accessor and drop
the now-unneeded density_matrix.h include from charge_mixing.h.
Timers move into the free functions with module_charge labels.
Add the direct parallel_orbitals.h include to esolver_gets.h, whose
value member previously relied on the removed transitive include.

* module_charge: decouple chg_dmr kernel from HContainer, mix raw buffers

Change module_charge::mix_dmr to take per-spin raw contiguous double
buffers and nnr instead of HContainer/DMR container references, and
drop the hcontainer.h include (and its atom_pair/parallel_orbitals
dependency chain) from chg_dmr.cpp. The sole call site in
esolver_ks_lcao.cpp now extracts the wrappers and saved buffers from
the DensityMatrix containers before calling the kernel. Move the
argument checks into a file-local check_dmr_inputs helper. The kernel
now depends only on the mixing module and MixingConfig.

* module_charge: refactor charge_mixing_rho free functions and cleanup

- Replace 17 PARAM.inp/globalv direct reads with cfg_ fields
- Unify mixing_tau: remove redundant member, use cfg_.mixing_tau
- Extract make_twobeta_mix as free function template in anonymous namespace
- Extract mix_tau_recip free function for kinetic energy density mixing
- Extract pack_rho_mag/unpack_rho_mag templates for nspin==2 dedup
- Hoist screen and inner_product lambdas before if-else chains (8+4 dups)
- Remove dead new_e_iteration member and its no-op if block
- Drop unused parameter.h include from charge_mixing_rho.cpp

* module_charge: split member functions into charge_mixing.cpp, free functions into chg_rho_detail.h

- Move mix_rho_recip/mix_rho_real/mix_rho from charge_mixing_rho.cpp to charge_mixing.cpp
- Create chg_rho_detail.h for make_twobeta_mix, pack_rho_mag, unpack_rho_mag templates and mix_tau_recip declaration
- charge_mixing_rho.cpp now only contains mix_tau_recip definition in module_charge::detail
- Restore accidentally deleted mix_uom member function

* module_charge: rename charge_{init,mixing_rho} to chg_{init,tau}, widen cube_io ofs_running to ostream

* charge_init.{cpp,h} -> chg_init.{cpp,h}: move Charge::init_rho stages
  (read_rho_from_file, init_rho_atomic_and_tau, load_rho_from_restart,
  init_rho_from_wfc) from Charge member functions to module_charge free
  functions, dropping the corresponding private declarations from
  charge.h. Continues the module_charge convention of stateless free
  functions in chg_* files.

* charge_mixing_rho.cpp -> chg_tau.cpp: rename for the module_charge
  short-underscore convention; the file only contains mix_tau_recip.

* Extract mix_tau_recip declaration from chg_rho_detail.h into a new
  chg_tau.h so chg_tau.cpp no longer pulls in the detail template
  helpers (make_twobeta_mix / pack_rho_mag / unpack_rho_mag).
  charge_mixing.cpp adds chg_tau.h while keeping chg_rho_detail.h for
  the template helpers it still uses.

* Widen ModuleIO::read_vdata_palgrid's ofs_running parameter from
  std::ofstream& to std::ostream& (cube_io.h / read_cube.cpp). The
  body only uses operator<<, so std::ostream& is sufficient; this
  fixes the chg_init.cpp compile error where read_rho_file /
  read_kin_file (per project rules, std::ostream&) could not bind to
  the old std::ofstream& parameter. Existing callers passing
  std::ofstream& (GlobalV::ofs_running, test fixture) convert
  implicitly via base-class reference.

Build lists updated: source/Makefile.Objects and
source/source_estate/{CMakeLists.txt,test/CMakeLists.txt}.

Verification: chg_init.* changes compile-verified by user before
this session; chg_tau rename and chg_tau.h extraction not yet
compile-verified; cube_io type widening not yet compile-verified.

* module_charge: rename charge_mixing.{h,cpp} to chg_mix.{h,cpp}, test to test_chg_mix.cpp

Pure rename, no logic change. Updates include guard, 12 #include sites,
CMakeLists (source_estate + test), and Makefile.Objects. CMake target
MODULE_ESTATE_charge_mixing kept (no external references). Class name
Charge_Mixing and module_charge namespace unchanged.

* module_charge: remove duplicate doc block comments (Phase 1a)

Remove or rephrase 14 duplicate comment lines across 7 files to
eliminate all duplicate_doc_block quality-score deductions.

- chg_mix.cpp: remove 7 duplicate comments in mix_rho_real that
  repeated mix_rho_recip's broyden/Kerker/magabs annotations
- chg_init.cpp: remove 2 duplicate comments in read_kin_file that
  repeated read_rho_file's binary-read and ParaWorld bridge notes
- chg_symm_detail.cpp: remove 1 duplicate step comment in psymmg_soc
- charge.h: rephrase kin_r_save comment to avoid repetition
- chg_extra.h: rephrase beta comment to avoid repetition
- chg_symm.cpp: remove 1 duplicate vector-management comment
- chg_precond.cpp: remove 1 duplicate Kerker comment

* module_charge: replace auto with explicit std::function types (Phase 1b)

Replace 14 auto-keyword lambda declarations with explicit
std::function types to eliminate all auto_keyword quality-score
deductions.

- chg_mix.cpp: 10 auto -> std::function (inner_product, screen,
  twobeta_mix in mix_rho_recip and mix_rho_real)
- chg_drho.cpp: 2 auto -> std::function<double()> (part_of_noncolin,
  part_of_rho)
- chg_tools.cpp: 1 auto -> std::function<void(int,int)> (kernel)
- chg_symm_detail.cpp: 1 auto -> std::function (build_wspin)

Added #include <functional> to all four files.

* module_charge: wrap lines over 120 chars (Phase 1c)

Break 21 lines exceeding the 120-char limit across 7 files to
eliminate all line_too_long quality-score deductions.

- charge.cpp: 3 WARNING_QUIT/cout lines split
- chg_atomic.cpp: 5 Simpson_Integral/exp/assert lines split
- chg_drho.cpp: 2 conj-product sum lines split
- chg_init.cpp: 1 warning message string split
- chg_mix.cpp: 5 make_twobeta_mix/recip_to_real/if_scf_oscillate lines split
- chg_mix.h: 3 member declaration/comment lines shortened
- chg_symm_detail.cpp: 2 MPI_Recv lines split

* module_charge: remove default parameter from Charge::init_rho (Phase 1d)

Remove the default nullptr values from init_rho's klist and wfcpw
parameters and update the two call sites (esolver_of.cpp,
esolver_double_xc.cpp) that relied on the defaults to pass nullptr
explicitly.

* module_charge: replace raw new/delete with std::vector and unique_ptr (Phase 2a-2d)

Replace all raw new/delete allocations in 4 files with RAII
containers to eliminate raw_new_keyword and unpaired_new_delete
quality-score deductions.

- chg_tools.cpp: 1 new -> std::vector<double> (aux buffer)
- chg_extra.cpp: 4 new -> std::vector<std::vector<double>> (rho_atom
  in extrapolate_charge and find_alpha_and_beta)
- chg_symm_detail.cpp: 14 new -> std::vector (rhog_piece, ig2isz,
  ipsz2ipw, nstnz_start, fftixy2is, rhogtot, ig2isztot, ixyz2ipw
  across reduce_to_fullrhog, rhog_piece_to_all, psymmg, psymmg_soc)
- chg_mix.{h,cpp}: 5 new + 5 unpaired -> std::unique_ptr for
  mixing and mixing_highf members; destructor and init_mixing
  simplified; get_mixing() returns .get()

charge.cpp (18 raw new) deferred to Phase 2e due to wider impact.

* module_charge: replace raw new/delete in Charge with vector-backed storage (Phase 2e)

Replace all 18 raw new and 10 unpaired delete in charge.cpp with
std::vector-backed storage to eliminate raw_new_keyword and
unpaired_new_delete deductions.

- charge.h: add _ptrs_rho, _ptrs_rhog, _ptrs_rho_save, _ptrs_rhog_save,
  _ptrs_kin_r, _ptrs_kin_r_save (std::vector<double*> / complex*),
  and _space_rho_core, _space_rhog_core (std::vector data buffers)
- charge.cpp allocate(): replace new double*[nspin] with vector resize;
  rho = _ptrs_rho.data() preserves double** interface
- charge.cpp init_final_scf(): replace both outer pointer and inner
  data new calls with _space_* vectors
- charge.cpp destroy(): replace delete[] with vector::clear() and
  nullptr assignment

charge.cpp score: 47 -> 69, now passing the 60 threshold.
Module average: 85.0 -> 85.7, 30/33 files passing.

* module_charge: replace std::make_unique with C++11-compatible unique_ptr(new T) (fix)

std::make_unique is a C++14 feature; the repo baseline is C++11.
Replace 4 make_unique calls with std::unique_ptr<T>(new T(...)) to
eliminate the post_cpp11_feature deduction (-40).

chg_mix.cpp score: 0 -> 15, module average: 85.7 -> 86.1.

* module_charge: fix duplicate doc block in charge.cpp init_final_scf

* module_charge: aggregate chgmixing_ks parameters into ScfMixingCtx struct (Phase 3a)

Replace 14-parameter chgmixing_ks with 7-parameter version by
grouping SCF convergence thresholds and status flags into a new
ScfMixingCtx struct, and deriving nrxx from chr.rhopw->nrxx.

- chg_routine.h: define ScfMixingCtx struct (hsolver_error, scf_thr,
  scf_ene_thr, converged_u, drho, oscillate_esolver, conv_esolver)
- chg_routine.cpp: unpack ctx members at function entry
- esolver_ks.cpp: pack ctx before call, unpack after

chg_routine.cpp score: 63 -> 70, too_many_parameters eliminated.

* module_charge: aggregate read_rho_file/read_kin_file parameters into ReadCfg (Phase 3b)

Replace 9-parameter read_rho_file and read_kin_file with 5-parameter
versions by grouping suffix, readin_dir, rank, ofs_running, ofs_warning
into a ReadCfg struct in the anonymous namespace.

chg_init.cpp score: 66 -> 70, too_many_parameters eliminated.

* module_charge: aggregate non_linear_core_correction parameters into NlcCtx (Phase 3c)

Replace 10-parameter non_linear_core_correction with 2-parameter
version by grouping all input data into a new NlcCtx struct.

chg_tools.cpp score: 96 -> 100, too_many_parameters eliminated.

* module_charge: split chg_mix.cpp into init and rho mixing files (Phase 4a)

Move mix_rho_recip, mix_rho_real, and mix_rho (440 lines) from
chg_mix.cpp into a new chg_mix_rho.cpp to eliminate file_too_long
deduction (-10).

- chg_mix.cpp: 727 -> 286 lines (constructor, set_mixing,
  init_mixing, set_rhopw, mix_reset, if_scf_oscillate,
  allocate_mixing_uom, mix_uom)
- chg_mix_rho.cpp: new file, 440 lines (mix_rho_recip,
  mix_rho_real, mix_rho)
- CMakeLists.txt: add chg_mix_rho.cpp to library and test targets

chg_mix.cpp score: 15 -> 60, now passing the 60 threshold.
32/34 files passing, module average improved.

* module_charge: split chg_drho.cpp and decompose inner product functions (Phase 4b)

Move inner_product_recip_rho and inner_product_recip_hartree from
chg_drho.cpp into a new chg_drho_inner.cpp, and decompose each
into per-nspin helper functions to reduce cyclomatic complexity.

- chg_drho.cpp: 520 -> 161 lines (cal_drho, cal_dkin,
  inner_product_real); score 49 -> 97
- chg_drho_inner.cpp: new file, 310 lines; score 100
  - inner_product_recip_rho decomposed into recip_rho_nspin1,
    recip_rho_nspin2, recip_rho_nspin4_mag helpers (CC 29 -> ~5 each)
  - inner_product_recip_hartree decomposed into
    recip_hartree_nspin2, recip_hartree_nspin4_trad,
    recip_hartree_nspin4_angle helpers (CC 37 -> ~5 each)
  - shared coulomb_sum_single extracted
- CMakeLists.txt: add chg_drho_inner.cpp to library and test targets

34/35 files passing, only chg_atomic.cpp remains below 60.

* refactor(module_charge): split atomic_rho and remove ZEROS in charge mixing

chg_atomic.cpp:
- Decompose atomic_rho (CC=60) into per-nspin helpers in
  chg_atomic_inner.cpp; CC reduced to 7, score 40->100.
- Replace all PARAM.inp.nelec/domag/domag_z/test_charge and
  GlobalV::ofs_warning with explicit AtomicRhoCfg parameter.
- Remove unused parameter.h include.
- Add chg_atomic_detail.h declaring detail helpers and RhoG3dCtx.

chg_init/chg_extra/esolver_*:
- Pass AtomicRhoCfg through call sites of atomic_rho,
  extrapolate_charge, and update_delta_rho.

Bug fixes:
- chg_drho_inner.cpp: fix duplicate const (const MixingConfig const&
  -> const MixingConfig&) and add detail:: prefix to helper calls.
- chg_mix_rho.cpp: use mixing.get()/mixing_highf.get() for unique_ptr.
- chg_tools.cpp: fix numeric -> numeric[it] in set_rho_core.

Memory safety / cleanup:
- Replace ModuleBase::GlobalFunc::ZEROS with std::fill in charge.cpp,
  chg_symm_detail.cpp, chg_tools.cpp; remove redundant ZEROS calls
  that precede full overwrites in chg_dmr.cpp and chg_mix_rho.cpp.

* Refactor: remove redundant Charge& overload of cal_rhog_symm_soc

The Charge& overload only forwarded chr.rho/chr.rhog to the raw-array
overload and had a single internal call site. Inline the member access
at that call site and drop the wrapper declaration and definition.

* module_charge: fix stale TITLE/timer labels and drop unused xc_functional.h includes

mix_tau_recip is now a free function in module_charge::detail, so update
its TITLE/timer labels from the legacy "Charge_Mixing" to "module_charge"
to match the convention of other free functions in the directory. Also
remove the unused xc_functional.h includes from chg_tau.cpp and
chg_symm_detail.cpp (label/include cleanup only, no behavior change).

* module_charge: remove redundant #ifdef __MPI guards around parallel wrappers

Parallel_Reduce::reduce_pool and Parallel_Common::bcast_double already
compile to no-op stubs when __MPI is undefined, so the outer guards add
nothing. Remove 11 such guards in chg_tools.cpp, chg_drho.cpp,
chg_drho_inner.cpp, chg_atomic_inner.cpp and chg_mix.cpp.

Guards enclosing raw MPI calls or MPI/serial dual paths are kept
(chg_parallel, chg_symm_detail, chg_routine BP_WORLD bcast, chg_extra.h).

* module_charge: decouple chg_routine from spin_constrain singleton

- forward-declare Plus_U_Base in chg_routine.h instead of including dftu_base.h
- query DeltaSpin mag_converged in ESolver_KS_PW and pass it to chgmixing_ks_pw

* module_charge: remove PARAM dependencies via explicit configuration structs

Remove the last four direct includes of parameter.h in module_charge
(chg_mix, chg_parallel, charge, chg_init) and the implicit PARAM.globalv.ks_run
read in chg_routine. INPUT values are now passed explicitly:

- MixingConfig gains scf_nmax for the drho oscillation history
- reduce_diff_pools/rho_mpi/kin_r_mpi take kpar, all_ks_run, bndpar, nspin,
  out_elf from callers instead of GlobalV::KPAR/PARAM
- Charge::kin_density/allocate/check_rho/renormalize_rho/init_final_scf take
  out_elf/test_charge/nelec as arguments with validation asserts
- new InitRhoCfg aggregates INPUT values for init_rho
- ScfMixingCtx gains ks_run; dm2rho takes nelec and drops its default
  skip_normalize argument per governance rule 5

No behavior change: save_rho_before_sum_band now uses the member nspin
set by allocate, identical to the previously read PARAM.inp.nspin.

* module_charge: restore #ifdef __MPI guards around parallel wrapper calls

The guards removed in 7a0013848 are load-bearing for serial-built unit
tests: source_estate/test strips __MPI from test translation units via
abacus_disable_feature_definitions, but links libbase built with __MPI,
whose explicit Parallel_Reduce instantiations contain real MPI calls.
Unguarded calls in the test TUs therefore bound to MPI_Allreduce and
abort with "called before MPI_INIT", failing MODULE_ESTATE_charge_test
and MODULE_ESTATE_charge_mixing.

Restore all 11 call-site guards in chg_tools.cpp, chg_atomic_inner.cpp,
chg_drho.cpp, chg_drho_inner.cpp and chg_mix.cpp. No behavior change for
MPI or serial production builds.

* Remove dead PAW compensation charge members

nhat, nhat_save in Charge and nhat_mdata in Charge_Mixing have had
no references since #6225 removed the PAW code; drop the orphaned
declarations and update the related comment.

* Refactor: remove unused Charge::prenspin member

prenspin recorded the spin-channel count read from legacy cube charge
files and drove collinear-to-noncollinear rearrangement in init_rho.
After read_rho was replaced by binary read_rhog (#5323, #5362) the
value is neither written nor read anywhere, so drop the dead member.

* Refactor: move Charge::cal_rho2ne/check_rho to module_charge free functions

- Add module_charge::check_rho in chg_tools.{h,cpp} with grid/geometry
  parameters passed explicitly; preserve all branches, thresholds and
  warning/abort messages of Charge::check_rho
- Remove the Charge::cal_rho2ne forwarding wrapper and Charge::check_rho
- Update the three esolver call sites (ks/of/double_xc) to pass rho,
  nspin, rhopw grid sizes and ucell.omega explicitly
- Drop the check_rho stubs in elecstate_pw/base tests and switch
  charge_test to the free cal_rho2ne
- Add test_chg_tools.cpp covering cal_rho2ne, total/spin-polarized
  checks, mismatch warning path and negative-channel aborts

* Refactor: remove redundant Charge::omega_ pointer

- Charge::sum_rho() now reads the cell volume from rhopw->omega, which
  is computed from the same lat0/latvec as ucell.omega and is already
  dereferenced on the same line for nxyz; this also makes the volume
  consistent with the grid rho lives on
- Drop the Charge::omega_ member, its set_omega() setter and the
  chg_init.cpp call site, removing a raw-pointer dependency on the
  UnitCell lifetime; update charge_test accordingly

Verified: MODULE_ESTATE_charge_test and MODULE_ESTATE_chg_tools pass,
elecstate library rebuilds cleanly.

* Remove dead Charge::init_final_scf and allocate_rho_final_scf

init_final_scf has had no production callers since the nscf refactor
(c6ae01236); its only remaining caller was the unit test added in
ba8b7ce9a. After the vector-backed storage refactor it was also a
broken duplicate of Charge::allocate: it never set nspin/nrxx/nxyz/
ngmc and skipped the kin_r buffers. Remove the function, its one-shot
guard flag, and the corresponding test case; destroy() now keys solely
on allocate_rho since vector storage self-manages cleanup.

* Refactor: pass rhopw explicitly to chg_init/chg_routine/chg_extra/chg_symm

Remove implicit reads of chr.rhopw/chr.ngmc from four module_charge files:
- chg_symm.cpp: size kin_g by the rho_basis used for its FFTs
- chg_routine: chgmixing_ks takes const PW_Basis&
- chg_init: orchestrator and four stage helpers take const PW_Basis&;
  the Charge::init_rho member signature is unchanged
- chg_extra: extrapolate_charge/update_delta_rho take const PW_Basis&

Call sites pass *chr.rhopw at the KS boundary or *pw_rhod where the
binding (esolver_fp.cpp chr.set_rhopw(pw_rhod)) makes them identical.
Verified: affected TUs compile and MODULE_ESTATE_charge_extra passes.

* Comments: add TODOs for LCAO+USPP double-grid follow-ups

Record the smooth/dense grid split to revisit if LCAO is ever allowed
with USPP: symmetrize_rho callers pass different grids, and the
ndx/ndy/ndz input path lacks the LCAO guard the ecutrho path has.

* Refactor: replace sticky Charge::cal_elf flag with explicit symm_kin argument

cal_elf was set to true once during ELF output and never reset, so every
later density symmetrization in the same run redundantly symmetrized
kin_r. Replace the mutable workflow flag with an explicit bool parameter
on the Charge& overload of module_charge::cal_rhog_symm:
- ctrl_output_fp passes true right before write_elf consumes kin_r
- symmetrize_rho wrapper and other callers pass XC_Functional::get_ked_flag()

Verified: full incremental build, read_wf2rho unit tests (serial/4 MPI),
write_elf logic test, and tests/01_PW/scf_out_elf (E difference 5e-10 eV,
ELF cube passes CompareFile.py at 3-decimal tolerance).

* Refactor: resolve mixing_tau at config assembly, drop XC dependency from chg_mix

esolver_ks now resolves mix_cfg.mixing_tau = inp.mixing_tau &&
XC_Functional::get_ked_flag() at the single production config assembly
point, so chg_mix/chg_mix_rho no longer query the XC global inside tau
mixing branches (6 sites). test_chg_mix mirrors the resolution in
make_cfg() and sets ked_flag before set_mixing where tau mixing is
expected. Also drop an unused xc_functional.h include from
chg_drho_inner.cpp.

Verified: full incremental build clean; MODULE_ESTATE_charge_mixing
11/11 tests pass; MODULE_ESTATE_charge/chg test suites all pass
(serial + 4-rank MPI).

* Fix: restore complete types in chg_drho_inner.cpp after include removal

Removing xc_functional.h in 87b818f4c broke compilation: the include was
load-bearing transitively, supplying the complete ModulePW::PW_Basis type
and ModuleBase::TITLE. Add the direct includes instead (pw_basis.h,
global_function.h) per IWYU.

Verified: make -j16 exits 0 with full log retained (previous verification
was invalid: a tail pipe masked both the exit code and the errors).

* Refactor: derive tau symmetrization/reduction from kin_r buffer existence

The Charge& cal_rhog_symm overload and rho_mpi/kin_r_mpi queried
XC_Functional::get_ked_flag() (plus a caller-supplied out_elf/symm_kin
flag) to decide whether to touch kin_r. Since Charge::allocate allocates
kin_r exactly when meta-GGA or ELF output needs it, both now check
chr.kin_r != nullptr directly, dropping the XC dependency and the extra
boolean parameters:
- rho_mpi/kin_r_mpi lose the out_elf parameter (2 production, 3 test
  call sites updated)
- the Charge& cal_rhog_symm overload loses the symm_kin parameter
  (ctrl_output_fp, setup_pot, read_wf2rho, update_state_rdmft revert to
  4 arguments); the raw-pointer overload now checks kin_r != nullptr
  only
- module_charge keeps XC references only in charge.cpp, chg_init.cpp,
  chg_drho.cpp (semantic "is meta-GGA" sites, resolved next)

Verified: make -j16 exit 0; 14/14 ctest charge/elecstate/read_wf2rho
tests (serial + 4-rank MPI); tests/01_PW/scf_out_elf integration case
reproduces the reference energy (-194.623411265 eV, diff 5e-10) and the
ELF cube passes CompareFile.py at 3-decimal tolerance.

* Refactor: remove module_xc dependency from module_charge (meta_gga state)

module_charge queried XC_Functional::get_ked_flag() at 5 semantic
"is meta-GGA" sites (tau TF init, tau file read, tau save, tau residual,
tau mixing resolution). Resolve the flag at upper layers instead:
- Charge::allocate takes an explicit meta_gga argument and stores it as
  object state; save_rho_before_sum_band and cal_dkin read it
- InitRhoCfg gains a meta_gga field, filled at the 3 esolver config
  assembly points (ks/of/double_xc)
- delete Charge::kin_density(); 6 esolver call sites inline
  get_ked_flag() || (out_elf[0] > 0) for buffer allocation and pass
  get_ked_flag() as meta_gga; non-SCF allocations pass false
- charge_test mirrors the inline expression

module_charge now has zero references to module_xc.

Verified: make -j16 exit 0 (full log); 14/14 charge/elecstate/
read_wf2rho ctests (serial + 4-rank MPI), including the mGGA tau mixing
and tau-save branches; tests/01_PW/scf_out_elf reproduces reference
energy (-194.623411265 eV, diff 5e-10) and the ELF cube passes
CompareFile.py at 3-decimal tolerance. A SCAN integration case
(205_PW_SCAN) still requires a libxc-enabled build/CI run.

* Fix: allow null rho buffers on ranks with empty real-space grid partition

pack_rho_mag/unpack_rho_mag in chg_rho_detail.h quit whenever any buffer
pointer is null. A rank may legitimately own zero real-space grid points
(nrxx == 0) when the grid is decomposed across more processes than it has
z-slabs (e.g. a 3x3x3 big-cell grid on 4 processes leaves one rank with
no slab); its zero-sized vectors then return null data() pointers even
though the packing loops perform no access. The unconditional check made
LCAO nspin==2 real-space mixing abort with "pack_rho_mag pointer is null"
on such ranks.

Restrict the null-pointer check to n > 0, matching the convention already
used by Parallel_Grid::reduce (only a null buffer with a non-zero size is
a genuine bug). n < 0 remains a hard error. Regression introduced in
d9685d4eb when the inline packing loops were extracted into these helpers.

* Refactor: move rhog_io into module_charge as chg_rhog_io

Relocate source_estate/rhog_io.{h,cpp} to source_estate/module_charge/
under the module_charge namespace, rename include guard to CHG_RHOG_IO_H,
and update the warning tags emitted at runtime. Update both callers
(chg_init.cpp, esolver_fp.cpp) and build files; adapt test_rhog_io.cpp in
place ahead of its move in a follow-up commit. No behavior change.

* Refactor: create module_charge/test with the rhog io unit test

Move test_rhog_io.cpp into module_charge/test/test_chg_rhog_io.cpp with
its support data charge-density.dat, register the new test subdirectory,
and rename the target to MODULE_CHARGE_rhog_io. Remove the migrated
AddTest block from the legacy source_estate/test/CMakeLists.txt.

* Refactor: move charge and charge-extra unit tests into module_charge/test

Rename charge_test.cpp to test_charge.cpp and charge_extra_test.cpp to
test_chg_extra.cpp per the test naming rule, move prepare_unitcell.h
alongside its only users, and register MODULE_CHARGE_charge /
MODULE_CHARGE_extra in the module_charge test CMakeLists. No test data
moves: prepare_unitcell.h only sets file-name strings at runtime, and
the extra test only writes cube files into ./support/.

* Refactor: move mix, parallel and tools unit tests into module_charge/test

Relocate test_chg_mix.cpp (fixing its relative includes), test_chg_parallel.cpp
and test_chg_tools.cpp into module_charge/test, register MODULE_CHARGE_tools /
MODULE_CHARGE_mix / MODULE_CHARGE_parallel with the 4-process mpirun test,
and drop the migrated blocks from the legacy source_estate/test CMakeLists.

* Refactor: rename module_charge test dir to unittests and wire CI for it

Rename source_estate/module_charge/test to unittests (relative CMake
paths are immune to the move). Sync the referencing points: the
add_subdirectory call, the coverage lcov filter (add '*/unittests/*' so
test sources stay excluded from the report), a dedicated Module_Charge
ctest step in test.yml with MODULE_CHARGE added to the catch-all -E
list to avoid double execution, and unittests/ added to the
code_quality_score.py SKIP_DIRS.

* Fix: pass ucell.omega to Charge::sum_rho/renormalize_rho to fix NPT stress

Root cause: commit 34b441e1c ("Refactor: remove redundant Charge::omega_
pointer") changed Charge::sum_rho() to read the cell volume from
rhopw->omega instead of ucell.omega. In variable-cell calculations (NPT),
pw_rho/pw_rhod are NOT rebuilt on cell change (only pw_wfc is), so
rhopw->omega keeps the initial cell volume while ucell.omega is updated
every MD step. The stale volume made sum_rho() return a wrong electron
count, which made renormalize_rho() scale rho by the wrong factor,
corrupting the stress (deviation ~0.002 in 095_PW_NPT) while the total
energy stayed near-correct (variational, second-order sensitive).

Fix: add an explicit omega parameter to Charge::sum_rho() and
renormalize_rho(); all call sites (init_scf, chg_routine, LCAO dm2rho path
through HSolverLCAO/dmToRho, RDMFT update_charge, OFDFT renormalize_psi)
now pass ucell.omega. This mirrors the existing check_rho(..., ucell.omega)
pattern.

Also mark three other rhopw->omega users with BUG(investigate) comments:
get_local_pp_energy, cal_delta_escf, and Makov-Payne correction. These are
pre-existing and were not changed by the refactor; they may have the same
stale-volume issue in NPT and should be investigated separately.

Bisected to 34b441e1c over the 20260916 module_charge refactor branch.

* Fix: add omega arg to remaining dm2rho call sites

Missed four LCAO_domain::dm2rho call sites in the previous commit:
- lcao_set.cpp init_chg_dm (skip_normalize=true, omega unused)
- esolver_dm2rho.cpp
- esolver_ks_lcao_tddft.cpp weight_dm_rho
- module_dm/init_dm.cpp

All now pass ucell.omega.

* Fix: restore HamiltHSMatrix hs declaration in cal_mw_from_lambda

Accidentally removed the line while editing the comment.

* Fix: close_kerker_gg0 actually disables Kerker; drop dead mixing_gg0 members

The chg_precond refactor (commit 6d127d517) made the Kerker kernels read
cfg_ (immutable INPUT snapshot) instead of Charge_Mixing members, but
close_kerker_gg0() kept writing the now-dead mixing_gg0/mixing_gg0_mag
members. As a result, the non-separate-loop EXX path in exx_lri_interface.hpp
silently failed to disable Kerker after convergence.

Fix: add a kerker_disabled_ flag on Charge_Mixing that the mix_rho_recip/
mix_rho_real screening lambdas short-circuit on. The flag lives on the
object, not in cfg_, so the immutable INPUT snapshot invariant is preserved.

Also drop the now-dead members mixing_gg0/mixing_gg0_mag/mixing_gg0_min/
mixing_angle/mixing_dmr and the get_mixing_gg0() getter; set_mixing/init_mixing
now read these from cfg_ directly. Add CloseKerkerGg0DisablesScreenReal
regression test that compares close_kerker_gg0() output against the
cfg.mixing_gg0=0 baseline and proves the flag is load-bearing.

* Fix: relax over-strict null-buffer asserts for empty grid partitions

reduce_diff_pools and Parallel_Grid::reduce_across_pools still forbade
null buffers unconditionally, contradicting the rule documented at
parallel_grid.cpp:355-360. A rank with nrxx == 0 may legitimately hold
a null rho/kin_r pointer; the MPI calls below use count 0 and ignore
the buffer. Align both call sites with the documented rule.

* Fix: relax over-strict null-buffer assert in ParaRgridWorld::reduce_across_pools

Same pattern as the previous fix: a rank with nrxx == 0 legitimately
holds a null buffer, and MPI_Allreduce with count 0 ignores it.
Align with the rule documented at parallel_grid.cpp:355-360.

* Fix: allow nnr == 0 in DMR mixing for empty MPI partitions

nnr is local to each MPI rank and may legitimately be zero when no
atom pairs survive the cutoff on that rank. The previous check
aborted DMR mixing for such distributions, whereas the historical
implementation allowed empty blocks. Relax the guard in
check_dmr_inputs() and init_mixing_dmr() to reject only negative
nnr, and require non-null DMR buffers only when nnr > 0, matching
the established nrxx == 0 convention in module_charge.

* Fix: split reciprocal rho copy from real-space |m| rescale in mix_rho_recip

The nspin==4 && mixing_angle>0 branch of mix_rho_recip mixed two
distinct operations in one loop bounded by npw, but rho_magabs is
sized nrxx (real-space) and the new |m| is written back by
recip2real into rho_magabs[0..nrxx-1]. Reading rho_magabs[npw+ig]
goes out of bounds once npw+ig >= nrxx (AddressSanitizer reproduces
with nrxx=125, npw=93) and the loop bound npw leaves the real-space
tail [npw, nrxx) of {mx,my,mz} unscaled. Split into two loops: the
reciprocal rho copy stays bounded by npw, the magnetization rescale
is bounded by nrxx and reads rho_magabs[ir].

* Refactor: remove unused Charge_Mixing::conserve_setting

conserve_setting() was introduced by 420f1ad00 (DeltaSpin feature
merge, 2026-06-15) but never wired up: no production caller, no
test reference, and the DeltaSpin module does not touch
Charge_Mixing. Drop the dead declaration per the project rule that
unused functions and their tests be removed.

* Refactor: drop dead Charge_Mixing::tpiba2 member

tpiba2 was declared in chg_mix.h but never assigned by set_mixing()
nor read anywhere in the module. Grep across the whole source tree
confirms all tpiba2 references are either ucell.tpiba2 (a separate
UnitCell member) or local variables in unrelated modules. The
Charge_Mixing class never computed or used its own tpiba2 pointer;
only tpiba is consumed by the stateless Kerker kernels via
mix_rho_recip/mix_rho_real. Remove the dead declaration.

* Refactor: route Charge_Mixing getters through cfg_

get_mixing_mode(), get_mixing_beta(), get_mixing_ndim() previously
returned the legacy mirror members that set_mixing() kept in sync
with cfg_ by hand. With cfg_ now treated as the immutable INPUT
snapshot, route the public getters through cfg_ directly so there
is a single source of truth for INPUT parameters. External callers
(esolver_ks_lcao, lcao_others, pw_others) are unaffected since
signatures are unchanged. The legacy members remain in place for
now; they are dropped in a later step after internal readers are
migrated.

* Refactor: init_mixing constructs Mixing from cfg_ not legacy mirrors

init_mixing() branched on this->mixing_mode and passed
this->mixing_ndim/mixing_beta to the Broyden/Pulay/Plain_Mixing
constructors. These legacy mirrors were kept in sync with cfg_
manually by set_mixing(). Route through cfg_ directly so cfg_
remains the single source of INPUT parameters. The Mixing objects
themselves still copy beta/ndim into their own members at
construction; that is a one-time snapshot and not a continuous
sync surface, so it is left untouched.

* Refactor: mix_rho_recip/mix_rho_real read mixing_beta from cfg_

Both mix_rho_recip and mix_rho_real built the twobeta_mix functor by
reading this->mixing_beta / this->mixing_beta_mag, which are legacy
mirrors that set_mixing() kept in sync with cfg_. Route the six
construction sites through cfg_.mixing_beta / cfg_.mixing_beta_mag
so cfg_ is the single source of INPUT parameters consumed by the
mixing logic. Behavior is unchanged since the mirrors and cfg_
hold identical values after set_mixing().

* Refactor: set_mixing stops mirroring cfg_ into legacy members

set_mixing() copied mixing_mode, mixing_beta, mixing_beta_mag,
mixing_ndim from cfg into legacy mirror members, then validation
and logging read from the mirrors. Now that all internal readers
(init_mixing, mix_rho_recip, mix_rho_real, getters) read from
cfg_, the mirror writes are dead work. Drop them and route
validation and log output through cfg_ directly. omega and tpiba
remain pointer members because they alias external runtime state
(cell volume, lattice constant) that changes across SCF iterations
and so do not belong in MixingConfig (an immutable INPUT snapshot).

* Refactor: drop legacy Charge_Mixing mirror members; cfg_ is single source

Drop mixing_mode, mixing_beta, mixing_beta_mag, mixing_ndim mirror
members. After the previous commits every internal reader (getters,
init_mixing, mix_rho_recip, mix_rho_real, set_mixing validation
and log output) routes through cfg_, so the mirrors are dead state
that set_mixing() no longer writes. cfg_ is now the single source
of truth for INPUT mixing parameters.

Update test_chg_mix.cpp accordingly: the two assertions that
reached directly into CMtest.mixing_beta_mag and CMtest.mixing_mode
now read CMtest.get_mixing_config().mixing_beta_mag and
CMtest.get_mixing_mode(), matching the public API used by the
other assertions in the same block. No production caller accessed
these members directly (esolver_ks_lcao, lcao_others, pw_others
all used the getters), so the change is test-only on the consumer
side.

* Refactor: drop NSDMI from MixingConfig to force explicit construction

The non-static data member initializers in MixingConfig provided
plausible-looking defaults (e.g. mixing_beta=0.8, mixing_mode=
"broyden") that silently masked forgotten fields when a new field
was added but not wired up at construction sites. With the
defaults removed, every construction site must use aggregate
initialization (or copy-assign from a fully-initialized instance),
and a missing field yields value-initialized (zero/empty) members
that are far more likely to trip a test than the old defaults.
Combined with -Wmissing-field-initializers promoted to error in
the next commits, adding a field to MixingConfig without updating
all aggregate-initialization sites becomes a compile error.

* Refactor: aggregate-init MixingConfig in esolver_ks with pragma guard

Convert the 17-line field-by-field assignment of mix_cfg into a
single aggregate initialization in declaration order. Wrap it in
#pragma GCC diagnostic error "-Wmissing-field-initializers" so
that adding a field to MixingConfig without updating this list
becomes a compile error rather than silently using a default.
Each initializer is annotated with the field name it corresponds
to, making the declaration-order dependency auditable at a glance.

* Refactor: aggregate-init MixingConfig in test_chg_mix with pragma guard

Convert make_cfg()'s 17-line field-by-field assignment into a
single aggregate initialization in declaration order, matching
the esolver-side change. Wrap in the same
#pragma GCC diagnostic error "-Wmissing-field-initializers" so
that adding a field to MixingConfig without updating the test
helper is also a compile error. Both construction sites (esolver
and test) now fail at compile time if a field is missing, closing
the maintenance gap where a new field could silently fall back to
a default value.

* Fix: fail-fast guards in Charge_Mixing and update chg_mix tests

Add validation to turn latent misuse (skipped set_rhopw/set_mixing)
into clear WARNING_QUIT errors instead of null dereference or heap
corruption:
- init_mixing rejects a null rhopw
- if_scf_oscillate checks scf_nmax > 0 and iteration range
- mix_rho validates chr/chr->rhopw and the grid pointers

Fix three chg_mix unit tests that read cfg_ before set_mixing, which
caused a SIGSEGV in SCFOscillationTest and assertion failures in the
two inner-product tests.

* test(module_charge): add unit tests for chg_uspp and chg_dmr

Add test_chg_uspp.cpp covering split_dgrid/merge_dgrid (normal split,
round-trip, nspin=1/2, empty high-frequency/smooth boundaries, and
input-validation abort paths).

Add test_chg_dmr.cpp covering init_mixing_dmr/mix_dmr (nspin=1/2/4
mixing with Plain_Mixing analytically verified, empty-partition null
buffer allowance, and input-validation abort paths).

Wire both targets into unittests/CMakeLists.txt.

* test(module_charge): add unit tests for chg_precond, chg_drho, chg_drho_inner, chg_mix_rho

- test_chg_precond.cpp: kerker_screen_recip/real (early return, nspin=1/2/4
  filter, nspin=4 with mixing_angle resize, real-space matches reciprocal).
- test_chg_drho.cpp: inner_product_real, cal_drho real-space path
  (nspin=1/2/4+domag_z), cal_dkin (meta_gga false/true).
- test_chg_drho_inner.cpp: inner_product_recip_rho and
  inner_product_recip_hartree for nspin=1 with a single G component,
  analytically verified against the Coulomb metric.
- test_chg_mix_rho.cpp: mix_rho abort paths (null chr/chr->rhopw, unset
  rhopw, double_grid without rhodpw) and real-space plain mixing value.

Wire all four targets into unittests/CMakeLists.txt.

* test(module_charge): add unit tests for chg_symm, chg_symm_detail, chg_atomic, chg_atomic_inner

- test_chg_symm.cpp: symmetrize_rho / cal_rhog_symm / cal_rhog_symm_soc
  no-op paths when symm_flag == 0, for nspin=1 and nspin=4.
- test_chg_symm_detail.cpp: psymmg and psymmg_soc idempotence on a
  manually built D_4 point group over a serial cubic PW_Basis.
- test_chg_atomic_inner.cpp: compute_rhoatm USPP direct-copy branch and
  NCPP integrate+scale-to-zv branch (Gaussian rho_at with known analytic
  integral); normalize_and_check renormalizes uniform density to nelec.
- test_chg_atomic.cpp: atomic_rho ntype==0 path (skips atom loop) and
  spin_number_need==3 abort path.

Wire all four targets into unittests/CMakeLists.txt.

* test(module_charge): add chg_tau/chg_routine/chg_init tests; drop spurious XC_Functional stubs

Fourth batch of module_charge unit tests:
- test_chg_tau.cpp: mix_tau_recip abort paths (null chr/grid/mixing, nspin<1,
  double_grid without high-f mixer) and non-double-grid plain mixing value.
- test_chg_routine.cpp: chgmixing_ks_pw/lcao iter==1 restart-step setup, and
  chgmixing_ks convergence branches (conv_esolver true / drho<hsolver_error
  skip mix_rho).
- test_chg_init.cpp: init_rho "wfc" with null wfcpw abort, and "atomic" with
  ntype==0 + meta_gga Thomas-Fermi tau initialization.
Wire all three targets into unittests/CMakeLists.txt.

Cleanup: remove the XC_Functional::func_type / ked_flag definitions from
test_chg_drho, test_chg_mix_rho, test_chg_symm, test_chg_atomic_inner,
test_chg_atomic, test_chg_tau, test_chg_routine, and test_chg_init. None of
the non-test sources compiled into these targets reference these statics
(charge.cpp, chg_*.cpp, and the linked base/cell_info/planewave_serial/
symmetry libraries are clean), so the definitions were pure dead weight.
Also drop the now-unneeded xc_functional.h include from test_chg_mix_rho.cpp
and correct the stub comments.

* test(module_charge): fix broken includes in unit tests

- include chg_atomic_detail.h instead of nonexistent chg_atomic_inner.h
  in test_chg_atomic_inner.cpp; add math_integral.h for Simpson_Integral
- include chg_drho.h in test_chg_drho_inner.cpp for
  module_charge::inner_product_recip_hartree
- include source_cell/magnetism.h in tests that define Magnetism stubs
  (test_chg_drho, test_chg_symm, test_chg_tau, test_chg_mix_rho)
- fix nonexistent source_charge/mixing includes in test_chg_tau.cpp to
  source_base/module_mixing

* test(module_charge): fix link/build issues; temporarily disable routine/init targets

- test_chg_tau: use Plain_Mixing(beta) ctor and init_mixing_data with
  complex type_size (old set_mixing_beta/init_mixing no longer exist)
- test_chg_symm_detail: add Magnetism stub required by cell_info's
  unitcell.cpp, matching other tests in this directory
- test_chg_routine: adapt to two-arg set_rhopw and tpiba from ucell
- disable MODULE_CHARGE_routine and MODULE_CHARGE_init targets with
  documented reasons: their transitive dependencies (Plus_U_Base,
  elecstate, source_io) are deeply coupled; to be resolved later

* fix bug

* format tool_quit

* remove a test due to WARNING_QUIT funcitno

* delete the support file charge-density.dat because unittests never need it

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
Co-authored-by: Xiaoyang Zhang <tsfxwbbzxy@163.com>

* tests: take three charge/DM tests off #define private public (#7998)

Redo of the work in #7988 and #7990, both of which were closed while the
charge density module was being restructured. That restructuring landed in
#7972 and already did most of the decoupling those PRs proposed: allocate(),
renormalize_rho() and sum_rho() now take their inputs explicitly, the mixing
parameters are aggregated in a MixingConfig, and chg_mix.cpp / chg_drho.cpp /
charge.cpp are free of global parameter reads. What was left was the
test-side access.

Production changes are additive only - no existing signature moves and no
line is deleted from any production header:

  Charge::get_allocate_rho()          - report whether allocate() has run
  Charge_Mixing::get_rho_mdata()      - mirror the existing get_dmr_mdata()
  Charge_Mixing::get_tau_mdata()
  Charge_Mixing::set_mixing_config()  - pair for the existing getter, for
                                        callers that must update the snapshot
                                        without rebuilding the mixing history
  XC_Functional::set_func_type()      - pair for get_func_type()
  XC_Functional::set_ked_flag()       - pair for get_ked_flag()

Test changes:

  test_dm_r_init   - two sites move to the already public get_DMR_save()
  test_charge      - the global parameter scratchpad becomes fixture state
                     (32 refs -> 0); PW_Basis setup goes through the public
                     initgrids/initparameters/setuptransform sequence instead
                     of the protected distribute_r()/distribute_g()
  test_chg_mix     - the scratchpad becomes a fixture-owned MixingConfig
                     (163 refs -> 0); the three blocks that hand-wired
                     Charge::_space_* now take their buffers from the fixture,
                     which owns them as vectors and points the public
                     rho/rhog/kin_r views at them with the same stride

No expected value or tolerance was changed.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Refactor: replace the JSON path walker with schema operations (#7994)

* Fix: Correct FFTW version detection and NCCL build summary (#7995)

* Fix: Correct FFTW version detection and NCCL build summary

* Read FFTW version directly from adjacent pkg-config metadata

* Fix: restore LibRI centered cell folding in get_Born_von_Karmen_cells

The previous replacement dropped LibRI's Array_Operator::operator%
mapping (c % n + 3*n/2) % n - n/2, shifting cell coordinates from
[-n/2, n/2) to [0, n). Callers using exact coordinate keys (e.g.
58_KP_LR_BSE reading (-1,-1,-1) from a (2,2,2) Rlist) failed with
"R coordinates are not in Rlist". Reintroduce the centered folding
in both the 1D and recursive overloads.

* Fix: propagate CUDA::nvToolsExt through base's link interface

For CUDA < 12.9, NVTX symbols (nvtxRangePushA/nvtxRangePop) live in
libnvToolsExt. Since __USE_NVTX is defined on the OBJECT library base
(which compiles timer.cpp), every consumer of base's object files needs
that library on its link line. Linking it only to the main executable
left unit tests that link base directly (e.g. MODULE_CELL_SYMMETRY_analysis)
with undefined NVTX references on CUDA 12.2 CI. Attach the dependency to
base as INTERFACE so it propagates to the executable and all test targets.

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
Co-authored-by: Xiaoyang Zhang <tsfxwbbzxy@163.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: SY Wang <uwsy1059@qq.com>
Co-authored-by: Taoni Bao <baotaoni@pku.edu.cn>
* Refactor: add n-free accessors to OccupationMatrix (step 1 of removing the n index)

The radial-channel index n of occ_[iat][l][n][spin] is always 0 in every
production path. Add n-free overloads of get/get_save/set/mat/mat_save that
forward to n=0, so call sites can migrate before the underlying storage
drops the n dimension.

* Refactor: drop the n index from the LCAO occupation-matrix accumulation (step 2)

The radial-channel loop in reduce_and_symmetrize_occ_k,
accumulate_occ_k_for_ik and process_occ_channel_gamma always skips n != 0,
so remove the loop and the n parameter of accumulate_occ_channel_k/gamma;
call sites switch to the new n-free OccupationMatrix accessors.

* Refactor: drop the n index from the LCAO potential and energy paths (step 3)

cal_pot_onsite's radial-channel loop always skips n != 0, so remove it and
switch iatlnmipol2iwt reads to channel 0. get_onsite_pot and the three
calc_energy_* helpers lose their N/n parameters; the Yukawa U/J lookups are
fixed to channel 0, matching every existing call site.

* Refactor: drop the n index from the PW path and ijr helpers (step 4)

* Refactor: drop the n storage dimension from OccupationMatrix (step 5)

* Refactor: remove the n loops from the DFT+U IO path (step 6)

* Refactor: drop the per-n U/J storage from YukawaScreening (step 8)

* Refactor: add corr_iwt accessor and migrate DFT+U lookup consumers

* Refactor: store corr_iwt for the correlated channel only and drop iatlnmipol2iwt

* Fix: use transN for onsite gemm in cal_force_gamma

The second gemm in cal_force_gamma was using transT (same as the first),
making it an identical computation that overwrote the first result.
The correct contraction for the onsite contribution is dS/dR * rho^N
(no transpose), matching the structure of cal_force_k which uses
rho^C for diag and rho^N for onsite.

rho_pot_onsite = DM * V_onsite is not symmetric, so rho^T != rho^N.

Added formula comments to both cal_force_gamma and cal_force_k.

* Refactor: merge op_legacy contributeHk specializations + compliance cleanup

- Merge three identical contributeHk template specializations into one
  generic implementation (94 -> 47 lines)
- Split comma-separated variable declarations (governance rule 8)
- Replace exit(0) with WARNING_QUIT for MPI safety
- Remove dead npol variables and commented-out code
- Unify timer labels to DFTU_LCAO across the module

* Refactor: templatize occ accumulation and narrow public header

- Merge accumulate_occ_channel_k/gamma into template accumulate_occ_channel<T>
- Merge accumulate_occ_k_for_ik into template accumulate_occ_for_ik<T>
- Rename reduce_and_symmetrize_occ_k to reduce_and_symmetrize_occ
- Move all internal helpers into anonymous namespace
- Move cal_occ_mat_k/gamma into DFTU_LCAO namespace
- Narrow public header from 108 to 61 lines (3 public functions)

* Refactor: shorten internal helper names in dftu_nao_occ.cpp

accumulate_occ_channel -> acc_channel
accumulate_occ_for_ik -> acc_for_ik
accumulate_occ_over_kstar -> acc_over_kstar
reduce_and_symmetrize_occ -> reduce_symm
process_occ_channel_gamma -> acc_channel_gamma

* Refactor: extract for_adj_pair skeleton and split cal_fs_nao_r_impl

- Add for_adj_pair() in dftu_nao_ijr.h as common pair-walk skeleton
- Migrate accumulate_hr_for_iat0 and compute_occ_from_dmr to for_adj_pair
- Split cal_fs_nao_r_impl (240 lines) into build_nlm, acc_fs_pairs,
  reduce_force, reduce_stress helpers (~100 lines each)

* Refactor: migrate module_dftu tests to unittests/ with test_<source> naming

Follow the module_charge convention:
- test/ -> unittests/, one test file per source file
- split dftu_core_test into test_dftu_nao_pots + test_dftu_nao_energy
- split dftu_operator_test into test_dftu_nao_op_legacy + test_dftu_nao_ijr
  + test_dftu_nao_for_r + test_dftu_nao_str_r
- rename dftu_lcao_test to test_dftu_nao_op
- rename test_dftu_nao_ijr to test_dftu_nao_ijr (already correct)
- unittests/CMakeLists.txt follows module_charge pattern with
  abacus_disable_feature_definitions and KEEP_FEATURE_DEFINITIONS __MPI
  for the two tests that call MPI_Init in main()
- add TODO comments on dftu_nao_occ/fs_k/fs_r about testability refactoring

* Test: add unit tests for get_linear_index and reduce_force/reduce_stress

- Extract reduce_force/reduce_stress from anonymous namespace into
  public reduce_force_impl/reduce_stress_impl in dftu_nao_fs_r.h,
  implemented in new dftu_nao_fs_reduce.cpp to keep link closure minimal.
- Add test_dftu_nao_folding.cpp covering get_linear_index row-major
  ("cg") and column-major ("scalapack_gvx") indexing.
- Add test_dftu_nao_fs_r.cpp covering reduce_force_impl nspin scaling
  and reduce_stress_impl weight + Voigt-to-3x3 rearrangement.
- Use minimal per-file mocks (UnitCell, Magnetism, SepPot, Sep_Cell,
  Parallel_Orbitals, Parallel_Reduce::reduce_all<double>) to avoid
  heavy link closures.
- Link dftu_nao_fs_reduce.cpp into MODULE_DFTU_op which consumes
  dftu_nao_fs_r.cpp.

Verified: ctest -R MODULE_DFTU 9/9 passed.

* Test: add unit tests for accumulate_diag_force/stress

- Extract accumulate_diag_force and accumulate_diag_stress from
  anonymous namespace in dftu_nao_fs_k.cpp into new header
  dftu_nao_fs_accum.h (templates must be header-only for external
  instantiation). accumulate_onsite_force stays in dftu_nao_fs_k.cpp
  because it needs Plus_U_Base's occupation-matrix lookup.
- Add test_dftu_nao_fs_accum.cpp covering double/complex diagonal
  force accumulation (atom attribution via iwt2iat) and stress
  accumulation with factor scaling.
- Link real parallel_orbitals.cpp and keep __MPI so
  Parallel_2D::set_serial works (mock constructor previously
  shadowed the real path, causing nrow=-1).

Verified: ctest -R MODULE_DFTU 10/10 passed.

* Fix: add dftu_nao_fs_reduce.o to Makefile.Objects

dftu_nao_fs_reduce.cpp defines reduce_force_impl/reduce_stress_impl but
was missing from OBJS_DFTU, causing undefined-reference link errors in
the Makefile build. The CMake build already lists the source.

* Test: add unit tests for dftu_nao_fs_reduce and dftu_nao_fs_k

Add two new unit test targets for module_dftu:

- MODULE_DFTU_fs_reduce (6 tests): covers reduce_force_impl spin
  scaling (nspin=1/2/4) and reduce_stress_impl Voigt-to-tensor
  conversion, symmetry, and lat0/omega weight.

- MODULE_DFTU_fs_k (2 tests): covers DftuFsEnv reference semantics
  (unchanged storage, external mutability). Uses heap-allocated
  dependencies without destruction to avoid BLACS/Grid_Driver
  linkage in the test binary.

dftu_nao_occ.cpp and dftu_nao_adj.cpp were evaluated but skipped:
their core functions are documented as hard to unit-test and depend
on hamilt::Hamilt, Grid_Driver::Find_atom, and TwoCenterIntegrator::snap,
making meaningful unit tests impractical without heavy refactoring.

Verified: ctest -R "MODULE_DFTU_fs_reduce|MODULE_DFTU_fs_k" -V
  8/8 tests passed.

* reduce unittest time

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
…nergies and Structure_Factor take their INPUT values explicitly (#8004)

* source_estate: give Charge_Extra an explicit test seam

test_chg_extra.cpp reached into eleven private members and one private
method of Charge_Extra, and staged nspin/chg_extrap in the private half of
the parameter singleton even though Init_CE() already takes both as
arguments.

Charge_Extra gains accessors over its extrapolation state - the step
bookkeeping, the displacement and delta_rho histories, and the alpha/beta
coefficients - plus find_alpha_and_beta_for_testing(), which runs the
private solver against whatever histories are currently held so its
solution can be checked on its own rather than only through a full
extrapolate_charge() step.

extrapolate_charge() also gains has_float_data, which it forwards to
Structure_Factor::setup(). That parameter arrives in a later commit on this
branch; taking it as an argument is what lets chg_extra.cpp stay free of
global reads instead of sourcing the flag itself.

The test's writes to the parameter singleton become fixture state
(24 refs -> 0). The global_out_dir write was dead: nothing in this target's
sources reads it.

No expected value was changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* source_estate: cal_energies takes its INPUT flags explicitly

ElecState::cal_energies() read four INPUT flags out of the global parameter
singleton to decide which optional energy terms contribute, which left
elecstate_energy_test.cpp with no way to drive them except writing the
private half of that singleton.

The flags are now passed one by one rather than bundled, so each call site
states exactly which behaviour it is selecting:

  cal_energies(type, imp_sol, sc_mag_switch, dft_plus_u, assume_isolated)

All nine call sites are updated. Eight of them already hold an Input_para
(this->inp_ in the esolvers, inp in chg_routine.cpp) and so add no global
references at all; only rdmft.cpp, which has no Input_para in scope, reads
PARAM.inp directly.

The six nspin reads are not threaded through: ElecState already owns the
Charge that carries nspin, and elecstate_pw.cpp reads this->charge->nspin
in nine places already, so the same member is used here.
elecstate_energy.cpp now has no global reads left.

In the test, nine of the twelve keys the fixture wrote were dead for this
target - makov_payne.cpp is the only linked source that still reads the
singleton, and its branch is never reached because assume_isolated stays
"none". The rest become fixture state.

No expected value was changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* source_pw: Structure_Factor::setup takes has_float_data explicitly

setup() read PARAM.globalv.has_float_data to decide whether to build the
single-precision eigts copies, so structure_factor_test.cpp had to write the
private half of the parameter singleton to exercise the float path.

setup() now takes the flag as an argument. Charge_Extra::extrapolate_charge()
forwards the value it was given in the previous commit, which keeps
chg_extra.cpp free of global reads - it has none today and this would
otherwise have put two back.

All call sites are updated, including the two mock definitions of
Structure_Factor::setup in psi_init_unit_test.cpp and test_chg_extra.cpp,
which have to move in lockstep with the real signature.

The test's twelve private reads of the eigts arrays needed no new interface:
the public get_eigts1_data<FPTYPE>() family already returns exactly those
pointers, so the reads move to <float> / <double> instantiations.

No expected value was changed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…about a third of their wall time (#8006)

* tests: let Autotest.sh run integration cases concurrently

Each integration suite is a single ctest test, and Autotest.sh walks its
cases in a serial loop. On the CI runner a case takes `-np 4` MPI ranks and
OMP_NUM_THREADS=2, i.e. 8 of the 16 available threads, so half the machine
idles while the remaining cases queue up. These cases are a few atoms
apiece, where a wider OpenMP team buys almost nothing; spending the same
thread budget on several cases side by side is far better.

Add `-j <n>`, or `-j auto` which fits as many whole `-np`-sized cases as
the core budget allows. The default stays 1, so nothing changes for a
caller that does not ask for concurrency.

Measured with `taskset -c 0-15` (15 usable CPUs) and OMP_NUM_THREADS=2
exported the way the CI workflow does:

    tests/01_PW (138 cases)   serial 267 s   -j auto (3) 100 s   -j 4 80 s
    tests/07_OFDFT            serial  49 s   -j auto (3)  19 s
    tests/03_NAO_multik       serial 145 s   -j auto (3)  53 s

Concurrency changes nothing observable. For all three suites test.sum is
byte-identical between the serial and the concurrent run, the summary
counters match (01_PW reports 4 failed / 2 fatal / 867 properties either
way), and the RUN/OK/WARNING/ERROR lines appear in the same order, because
each case's console output is buffered and replayed in cases-file order.

Points worth recording:

- `nproc` honours $OMP_NUM_THREADS, which the CI workflow exports, so it
  reports the per-case thread count rather than the machine size. The core
  budget is therefore read with both OpenMP variables cleared.
- A concurrent run recomputes the per-case thread count instead of
  inheriting an $OMP_NUM_THREADS that was sized for one case at a time;
  `-o` still pins it explicitly.
- Address Sanitizer runs are forced back to -j 1, because every case
  appends to one shared diagnostics report.
- A case now runs in a subshell in both modes and its counters are
  aggregated from per-case files instead of shell state. That also removes
  a latent double count: the serial loop used to leave the last case's
  counters in scope, which added one spurious test.sum line.
- MPI binding is left as it is. In a harness that ran the cases without
  checking them, `--bind-to none` measured slower than the default binding
  (80.3 s against 65.2 s at 4 concurrent cases), so no launcher flag is
  added.

check_out is unchanged apart from two trailing-whitespace-only lines.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* tests: run the CPU integration suites with -j auto

Wire the concurrency from the previous commit into the suites that ctest
registers. `auto` divides the cores available to the process by the suite's
`-np`, so a 16-core CI runner takes 4 cases at a time while a small
developer machine stays serial.

The GPU suites keep their serial invocation: they share one device, and
concurrent cases would contend for its memory. The AddressSanitizer
variants are left alone as well; Autotest.sh forces those back to -j 1
regardless, since they share one diagnostics report.

Verified by configuring with -DBUILD_TESTING=ON and reading the registered
commands back out of the generated CTestTestfile.cmake: the 11 CPU suites
present in that configuration carry "-n" "4" "-j" "auto", and the GPU suite
does not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Zhang Zhili <185902695+zzlinpku@users.noreply.github.com>
* module_charge: normalize indentation and brace single-statement control flow

Mechanical cleanup as the first step of the module_charge governance
refactor: convert leading tabs to 4-space indentation (1011 occurrences
across 11 files) and add braces around all single-statement if/for/while
bodies (11 sites). No functional change.

* module_charge: aggregate Charge_Mixing params into MixingConfig

Introduce a MixingConfig POD that bundles the INPUT mixing parameters
with the runtime globals (nspin, scf_thr_type, double_grid), and change
set_mixing from a 12-argument interface to set_mixing(const MixingConfig&,
double&, double&). Charge_Mixing now stores the config and reads nspin /
scf_thr_type / double_grid from it instead of PARAM.inp / PARAM.globalv,
removing the direct PARAM reads in set_mixing and init_mixing.

The single production call site (esolver_ks.cpp) fills the config, and
the unit test drives set_mixing via a make_cfg() helper. The
'#define private public' access hack is kept for now with a TODO: the
test still must write Parameter::input/sys, Charge::_space_* and
XC_Functional privates, which need the Step 4/5 global-state
parameterization before it can be removed.

Verified: make -j30 MODULE_ESTATE_charge_mixing (build_max_para_test)
passes with no errors.

* module_charge: deduplicate twobeta_mix lambdas and replace raw new with std::vector

Extract the repeated two-beta mixing functor in mix_rho_recip/mix_rho_real
into a make_twobeta_mix<T> template helper (6 lambda copies removed), and
convert all local raw new[]/delete[] buffers in charge_mixing_rho.cpp to
zero-initialized std::vector, dropping the paired ZEROS calls.

* module_charge: move residual/inner-product globals into MixingConfig

Extend MixingConfig with gamma_only_pw/domag/domag_z so mix_resid.cpp
(get_drho, get_dkin, inner_product_recip_{rho,simple,hartree,real}) no
longer reads PARAM/GlobalV; all branches now consume this->cfg_.
inner_product_recip_rho's raw pointer-array views are switched to
std::vector. Production fills the three new fields in esolver_ks, and
the test fixture gains a sync_cfg() helper to push PARAM mutations into
cfg_ for the inner-product branch tests.

* module_charge: own Charge's _space_* storage with std::vector (Step 5a)

Replace the six private raw _space_rho/_space_rho_save/_space_rhog/
_space_rhog_save/_space_kin_r/_space_kin_r_save buffers with
std::vector, so Charge's underlying contiguous storage self-manages and
the matching delete[] calls in destroy() (which relied on reading
possibly-uninitialized pointers) go away. The public rho/rhog/rho_save/
rhog_save/kin_r/kin_r_save views keep their double**/complex** shape and
still alias the vector memory via .data(), so all external consumers are
unaffected. Tests that drove _space_* directly are adapted to
resize()/.data() and drop their manual delete[] of the buffers.

* module_charge: route chgmixing_ks through its inp parameter

chgmixing_ks already takes a const Input_para& inp but still read
PARAM.inp.mixing_restart / PARAM.inp.scf_nmax from the global. Use the
inp argument instead so the function no longer reads INPUT state through
the global for these two fields. PARAM.globalv.ks_run is a runtime
per-process flag (set from band-parallel topology), not an input, so it
is intentionally left as-is rather than threading it through the
interface.

* module_charge: split Charge::init_rho into per-stage private methods

init_rho had a cyclomatic complexity of 36 from five sequential stages
(file read, atomic fallback, Thomas-Fermi tau, restart load, wfc read)
interleaved through shared read_error/read_kin_error flags. Extract the
four branches into private methods -- read_rho_from_file,
init_rho_atomic_and_tau, load_rho_from_restart, init_rho_from_wfc -- and
leave init_rho as a thin sequence of stage calls. Logic is unchanged; the
error flags are threaded through as parameters. The deepest stage
(read_rho_from_file) now sits at complexity 19, down from 36 for the
monolith. The remaining global reads inside the stages are untouched and
deferred to a later parameterization step.

* module_charge: extract Charge density math into charge_math free functions

sum_rho, cal_rho2ne and non_linear_core_correction each used Charge
members only to reach a handful of scalars (nrxx/nxyz/omega) or the
reciprocal-shell table (gg_uniq/ngg); the rest of each body is pure
numerics. Move the three bodies into a new charge_math namespace as free
functions with those values passed explicitly, and leave the Charge
members as thin forwarding wrappers so no caller outside the module
changes. The kernels are now unit-testable in isolation and no longer
coupled to Charge state. One behavior note: the pre-quit debug line that
printed sum_rho to ofs_warning is dropped so the free function stays free
of global-stream dependencies. charge_math.cpp is wired into the estate
library and the charge_test target.

* module_charge: register charge_math.o in the hand-written Makefile build

The CMake build already picks up charge_math.cpp; mirror that in
Makefile.Objects so the legacy Makefile flow links the new charge_math
kernels too. The module_charge directory is already on VPATH, so adding
charge_math.o to the object list is sufficient.

* module_charge: extract Charge::atomic_rho into charge_atomic free function

Remove Charge::atomic_rho entirely and replace all call sites with
module_charge::atomic_rho(..., rhopw), eliminating the need for a thin
wrapper on the Charge class. This decouples atomic density initialization
from Charge's state and improves charge.cpp quality score from 2 to 44.

* module_charge: forbid Charge copies and guard tau.cube write

scf_out_chg_tau aborted in Parallel_Grid::reduce on
assert(rhoin != nullptr) because the kin_r_save[is] handed to
write_vdata_palgrid was not a valid buffer. After the _space_* storage
became std::vector (ecf5084d4), a copied/moved Charge leaves its
rho/kin_r views dangling into another object's vector buffer, and a
kin_r_save never allocated (ked_flag set after allocate) stays nullptr;
both surface as a null rhoin deep inside MPI gather instead of at the
source.

Delete Charge's copy constructor/assignment so any value copy of the
vector-aliasing views fails at compile time, and check kin_r_save in
ctrl_output_fp before writing tau.cube so a missing allocation reports
a clear message instead of tripping the MPI assert.

Verification: not run locally (per user request, user compiles).

* module_base: tolerate null grid buffer when a rank owns no grid points

scf_out_chg_tau (LCAO, SCAN, out_chg=1, 4 MPI ranks) aborted in
Parallel_Grid::reduce on assert(rhoin != nullptr). Bisecting between
83eb5d0f3 (good) and ecf5084d4 (bad) isolated the regression to
ecf5084d4, which moved Charge's _space_* storage from raw new[] to
std::vector.

Root cause: with 4 ranks the FFT grid is slab-decomposed so that the
last rank owns zero real-space points (nrxx == 0, confirmed via a
temporary diagnostic printing fn/is/rank/nrxx at the reduce call site).
Before ecf5084d4, _space_rho = new double[nspin * 0] == new double[0]
returned a unique non-null pointer, so rho_save[is] was non-null and the
assert passed. After the change, an empty vector's .data() returns
nullptr, so the rank with nrxx == 0 handed a null rhoin to reduce and
tripped the assert (Debug) or fed MPI_Gatherv a null buffer (Release).

A rank with nrxx == 0 is legitimate: MPI_Gatherv is invoked with
sendcount 0 and ignores the send buffer. Relax the assert to only flag a
null buffer when nrxx != 0, and revert the now-unneeded kin_r_save guard
in ctrl_output_fp (it would have falsely aborted on the nrxx == 0 rank).

Verification: Release build (build_max_para_test), ran
  cd tests/03_NAO_multik/scf_out_chg_tau &&
  OMP_NUM_THREADS=1 mpirun -np 4 ../../../build_max_para_test/abacus_max_para
Result: exit 0, chg.cube and tau.cube written; numerical comparison
against chg.cube.ref/tau.cube.ref gives maxdiff 0 (chg) and 1e-14 (tau).

* module_charge: extract Charge::set_rho_core into charge_math free function

Move set_rho_core to charge_math::set_rho_core with rho_core,
rhog_core and rhopw passed explicitly instead of reading Charge
state, and call charge_math::non_linear_core_correction directly.
Remove the now-unused Charge::non_linear_core_correction wrapper,
use std::vector for the rhocg/vg scratch buffers, update the
init_scf call site, and drop the obsolete member stubs in the
elecstate unit tests.

* module_charge: vectorize Charge_Extra history arrays and forbid copies

Replace the raw new[]/delete[] displacement arrays (dis_old1, dis_old2,
dis_now) with std::vector and remove the hand-written destructor. This
fixes a read of uninitialized pot_order when an object is destroyed
before Init_CE, a memory leak when Init_CE is called repeatedly, and a
double-free risk from the implicitly generated shallow copy. The copy
constructor and copy assignment are deleted so the molecular-dynamics
trajectory history cannot be silently forked. The unit test now checks
vector sizes instead of non-null pointers.

* Rename charge_math to chg_tools and unify namespace module_charge

- Rename module_charge/charge_math.{h,cpp} to chg_tools.{h,cpp} via git mv
- Change namespace charge_math to module_charge to match charge_atomic
  and chgmixing in the same directory
- Update include guard CHG_TOOLS_H and TITLE/timer labels accordingly
- Update call sites in init_scf.cpp, charge.cpp, charge_init.cpp
- Update build references in Makefile.Objects and both CMakeLists.txt

* module_charge: refactor Symmetry_rho class to free functions

Convert the stateless class Symmetry_rho into namespace module_charge
free functions and rename files for consistency:
  symm_rho.{h,cpp}      -> chg_symm.{h,cpp}
  symm_rho_detail.h     -> chg_symm_detail.h
  symm_rhog.cpp         -> chg_symm_detail.cpp

- 5 public functions become module_charge::symmetrize_rho / cal_rhog_symm
  (2 overloads) / cal_rhog_symm_soc (2 overloads)
- 2 cross-TU helpers (psymmg/psymmg_soc) moved to module_charge::detail
  via chg_symm_detail.h
- 3 internal MPI helpers moved to anonymous namespace
- Delete dead code psymm (real-space symmetrization, never called)
- Remove empty ctor/dtor and parallel_grid.h include
- Rename begin/begin_soc to cal_rhog_symm/cal_rhog_symm_soc for clarity
- Update timer/TITLE labels from "Symmetry_rho" to "module_charge"
- Migrate all 14 call sites and 1 test stub
- Remove obsolete Makefile special rule (no more name collision)

* module_charge: extract MixingConfig header and drop unused inner_product_recip_simple

Move MixingConfig from charge_mixing.h into its own mixing_config.h so
stateless residual kernels can include the config without dragging in
Charge_Mixing. Remove inner_product_recip_simple, which had no production
call sites, together with its unit test.

* module_gint: move gint_prec_ctrl from module_charge

Relocate gint_prec_ctrl.{h,cpp} and its test into module_gint, update the
include in esolver_ks_lcao.h and rewire the CMake/Makefile object lists.

* module_charge: extract mixing inner products into chg_drho free functions

Rename mix_resid.cpp to chg_drho.cpp and turn inner_product_real and
inner_product_recip_hartree into module_charge free functions declared
in chg_drho.h; inner_product_recip_rho, which is only shared with the
unit test, moves to module_charge::detail in chg_drho_detail.h.
Charge_Mixing loses the three private inner-product members and
mix_rho_recip/mix_rho_real bind the free functions through lambdas.
get_drho/get_dkin stay as members for this step.

* module_charge: hide cal_drho/cal_dkin in an anonymous namespace

Move the get_drho/get_dkin implementations into file-local cal_drho/
cal_dkin free functions with all inputs explicit; the public
Charge_Mixing methods become thin forwarding wrappers so esolver call
sites stay unchanged.

* module_gint: fix include path in test_gint_prec_ctrl after relocation

* module_charge: extract Kerker screen kernels into chg_precond free functions

Move Charge_Mixing::Kerker_screen_recip/real to module_charge namespace
as free functions in chg_precond.{h,cpp}, renaming mix_precond.cpp via
git mv. Config/grid/geometry are passed explicitly via MixingConfig,
PW_Basis*, and tpiba, eliminating the function's direct read of
PARAM.inp.nspin. Replace 8 std::bind call sites in charge_mixing_rho.cpp
with lambdas, update 2 commented-out bind sites in charge_mixing_dmr.cpp,
and rewrite 12 test call sites in charge_mixing_test.cpp to construct an
independent MixingConfig instead of poking at Charge_Mixing privates.
Drop the now-unused member function declarations from charge_mixing.h.

* module_charge: fix Makefile.Objects after mix_precond -> chg_precond rename

Update the non-CMake object list to track the renamed translation unit so
make-based builds do not reference the deleted mix_precond.o.

* module_charge: drop Charge_Mixing::get_drho/get_dkin wrappers

Expose cal_drho/cal_dkin as module_charge free functions in chg_drho.h
and let ESolver_KS call them directly with explicit arguments; add
Charge_Mixing::get_mixing_config() as a const observer for the config.

* module_charge: rename chgmixing.h/cpp to chg_routine.h/cpp

Align with the chg_<feature> naming pattern used in the same directory
(chg_drho, chg_precond, chg_symm, chg_tools). Update include guard to
CHG_ROUTINE_H, the self-include in chg_routine.cpp, the entry in
source_estate/CMakeLists.txt and source/Makefile.Objects, and the three
#include sites in esolver_ks{,_pw,_lcao}.cpp. Function names
(chgmixing_ks{,_pw,_lcao}) and TITLE/timer tags are intentionally left
unchanged to keep the diff minimal.

* module_charge: rename mixing_config.h to chg_mix_cfg.h

Rename the MixingConfig header to align with the chg_* naming
convention in module_charge. Update the include guard and the four
in-tree includers; no CMake change is needed since the header is not
listed explicitly.

* module_charge: convert Charge MPI helpers into chg_parallel free functions

Rename charge_mpi.cpp to chg_parallel.cpp and add chg_parallel.h, moving
the three stateless Charge member functions (reduce_diff_pools, rho_mpi,
kin_r_mpi) to module_charge namespace free functions that take the
Charge object explicitly. Remove their declarations from charge.h and
update all call sites in elecstate_pw, stress_mgga, read_wf2rho_pw and
sto_iter. Rename the unit test to test_chg_parallel.cpp and update the
test target name accordingly.

GlobalV/PARAM reads and the direct MPI_Allreduce in reduce_diff_pools
are preserved as pre-existing technical debt (migration-neutral).

* Rename charge_atomic files to chg_atomic

- Rename module_charge/charge_atomic.{h,cpp} to chg_atomic.{h,cpp}
- Update include guard to CHG_ATOMIC_H
- Update includes in charge_init.cpp and charge_extra.cpp
- Update source paths in CMakeLists.txt, test CMakeLists.txt
- Fix stale object names in Makefile.Objects: replace
  symm_rho_charge.o/symm_rhog.o with chg_symm.o/chg_symm_detail.o

* module_charge: extract USPP double-grid split/merge into chg_uspp free functions

Introduce module_charge::split_dgrid / merge_dgrid in chg_uspp.{h,cpp} as
RAII, parameter-explicit replacements for Charge_Mixing::divide_data /
combine_data / clean_data, which paired raw new[] with manual delete[]
across ~160 lines of mixing code.

- chg_uspp.{h,cpp}: stateless free functions in module_charge namespace;
  outputs are caller-pre-sized std::vector, no new/delete; parameter
  validation via WARNING_QUIT; TITLE/timer tags preserved
- charge_mixing_rho.cpp: rho and tau double-grid paths switched to the new
  functions; raw pointer aliases kept for !double_grid so the existing
  mixing call sites (nspin==1/2/4) are untouched
- CMakeLists.txt (source + test): wire chg_uspp.cpp

The legacy divide_data/combine_data/clean_data members are not yet removed;
that follows in a later step after the test is updated.

* module_charge: rewrite MixDivCombTest for the new split_dgrid/merge_dgrid

Drop the legacy alias-pointer assertions (EXPECT_EQ(datas, data.data()),
EXPECT_EQ(datas, nullptr) after clean_data) that coupled the test to the
old new[]/delete[] ownership model.

The rewritten case verifies the actual contract:
- split_dgrid fills smooth and high-frequency buffers with the dense
  data verbatim (per-element comparison)
- merge_dgrid is a left-inverse of split_dgrid (output == input)
- no explicit cleanup call is required: std::vector manages storage

Covers nspin == 1 and nspin == 2 paths.

* module_charge: drop legacy divide_data/combine_data/clean_data members

With the new module_charge::split_dgrid/merge_dgrid in chg_uspp.{h,cpp}
and all call sites in charge_mixing_rho.cpp migrated, the original
Charge_Mixing::divide_data / combine_data / clean_data members are dead.

- delete charge_mixing_uspp.cpp (the raw new[]/delete[] implementation)
- drop the three member declarations from charge_mixing.h
- remove charge_mixing_uspp.cpp from source/test CMakeLists.txt
- Makefile.Objects: drop charge_mixing_uspp.o, add chg_uspp.o
- refresh one stale comment in charge_mixing_rho.cpp to reference
  merge_dgrid instead of the removed combine_data

* module_charge: rename charge_extra files to chg_extra and move class into namespace

Rename charge_extra.h/cpp to chg_extra.h/cpp and wrap the Charge_Extra
class in the module_charge namespace, matching the rest of module_charge
(chg_atomic, chg_symm, chg_uspp). Update include guards, call sites in
esolver_fp.h and the unit test, and CMake/Makefile source lists.

* module_charge: extract DMR mixing into chg_dmr free functions

Move the DMR allocation/mixing logic out of Charge_Mixing members into
stateless module_charge functions (init_mixing_dmr, template mix_dmr
with explicit instantiation), passing the Mixing object, mixing data
and MixingConfig explicitly instead of reading PARAM. Merge the two
identical real/complex mix_dmr overloads, replace raw new[]/delete[]
of the magnetic buffers with std::vector, and de-duplicate the
two-beta mixing lambda into a file-local helper. The members stay as
thin timer-wrapped wrappers so external call sites are unchanged.

* module_charge: remove Charge_Mixing DMR wrappers, call chg_dmr directly

Delete charge_mixing_dmr.cpp and have the two call sites
(chg_routine.cpp, esolver_ks_lcao.cpp) invoke module_charge::
init_mixing_dmr/mix_dmr directly with the Mixing object, mixing data
and MixingConfig obtained through Charge_Mixing accessors. Expose the
owned DMR mixing history via a new get_dmr_mdata() accessor and drop
the now-unneeded density_matrix.h include from charge_mixing.h.
Timers move into the free functions with module_charge labels.
Add the direct parallel_orbitals.h include to esolver_gets.h, whose
value member previously relied on the removed transitive include.

* module_charge: decouple chg_dmr kernel from HContainer, mix raw buffers

Change module_charge::mix_dmr to take per-spin raw contiguous double
buffers and nnr instead of HContainer/DMR container references, and
drop the hcontainer.h include (and its atom_pair/parallel_orbitals
dependency chain) from chg_dmr.cpp. The sole call site in
esolver_ks_lcao.cpp now extracts the wrappers and saved buffers from
the DensityMatrix containers before calling the kernel. Move the
argument checks into a file-local check_dmr_inputs helper. The kernel
now depends only on the mixing module and MixingConfig.

* module_charge: refactor charge_mixing_rho free functions and cleanup

- Replace 17 PARAM.inp/globalv direct reads with cfg_ fields
- Unify mixing_tau: remove redundant member, use cfg_.mixing_tau
- Extract make_twobeta_mix as free function template in anonymous namespace
- Extract mix_tau_recip free function for kinetic energy density mixing
- Extract pack_rho_mag/unpack_rho_mag templates for nspin==2 dedup
- Hoist screen and inner_product lambdas before if-else chains (8+4 dups)
- Remove dead new_e_iteration member and its no-op if block
- Drop unused parameter.h include from charge_mixing_rho.cpp

* module_charge: split member functions into charge_mixing.cpp, free functions into chg_rho_detail.h

- Move mix_rho_recip/mix_rho_real/mix_rho from charge_mixing_rho.cpp to charge_mixing.cpp
- Create chg_rho_detail.h for make_twobeta_mix, pack_rho_mag, unpack_rho_mag templates and mix_tau_recip declaration
- charge_mixing_rho.cpp now only contains mix_tau_recip definition in module_charge::detail
- Restore accidentally deleted mix_uom member function

* module_charge: rename charge_{init,mixing_rho} to chg_{init,tau}, widen cube_io ofs_running to ostream

* charge_init.{cpp,h} -> chg_init.{cpp,h}: move Charge::init_rho stages
  (read_rho_from_file, init_rho_atomic_and_tau, load_rho_from_restart,
  init_rho_from_wfc) from Charge member functions to module_charge free
  functions, dropping the corresponding private declarations from
  charge.h. Continues the module_charge convention of stateless free
  functions in chg_* files.

* charge_mixing_rho.cpp -> chg_tau.cpp: rename for the module_charge
  short-underscore convention; the file only contains mix_tau_recip.

* Extract mix_tau_recip declaration from chg_rho_detail.h into a new
  chg_tau.h so chg_tau.cpp no longer pulls in the detail template
  helpers (make_twobeta_mix / pack_rho_mag / unpack_rho_mag).
  charge_mixing.cpp adds chg_tau.h while keeping chg_rho_detail.h for
  the template helpers it still uses.

* Widen ModuleIO::read_vdata_palgrid's ofs_running parameter from
  std::ofstream& to std::ostream& (cube_io.h / read_cube.cpp). The
  body only uses operator<<, so std::ostream& is sufficient; this
  fixes the chg_init.cpp compile error where read_rho_file /
  read_kin_file (per project rules, std::ostream&) could not bind to
  the old std::ofstream& parameter. Existing callers passing
  std::ofstream& (GlobalV::ofs_running, test fixture) convert
  implicitly via base-class reference.

Build lists updated: source/Makefile.Objects and
source/source_estate/{CMakeLists.txt,test/CMakeLists.txt}.

Verification: chg_init.* changes compile-verified by user before
this session; chg_tau rename and chg_tau.h extraction not yet
compile-verified; cube_io type widening not yet compile-verified.

* module_charge: rename charge_mixing.{h,cpp} to chg_mix.{h,cpp}, test to test_chg_mix.cpp

Pure rename, no logic change. Updates include guard, 12 #include sites,
CMakeLists (source_estate + test), and Makefile.Objects. CMake target
MODULE_ESTATE_charge_mixing kept (no external references). Class name
Charge_Mixing and module_charge namespace unchanged.

* module_charge: remove duplicate doc block comments (Phase 1a)

Remove or rephrase 14 duplicate comment lines across 7 files to
eliminate all duplicate_doc_block quality-score deductions.

- chg_mix.cpp: remove 7 duplicate comments in mix_rho_real that
  repeated mix_rho_recip's broyden/Kerker/magabs annotations
- chg_init.cpp: remove 2 duplicate comments in read_kin_file that
  repeated read_rho_file's binary-read and ParaWorld bridge notes
- chg_symm_detail.cpp: remove 1 duplicate step comment in psymmg_soc
- charge.h: rephrase kin_r_save comment to avoid repetition
- chg_extra.h: rephrase beta comment to avoid repetition
- chg_symm.cpp: remove 1 duplicate vector-management comment
- chg_precond.cpp: remove 1 duplicate Kerker comment

* module_charge: replace auto with explicit std::function types (Phase 1b)

Replace 14 auto-keyword lambda declarations with explicit
std::function types to eliminate all auto_keyword quality-score
deductions.

- chg_mix.cpp: 10 auto -> std::function (inner_product, screen,
  twobeta_mix in mix_rho_recip and mix_rho_real)
- chg_drho.cpp: 2 auto -> std::function<double()> (part_of_noncolin,
  part_of_rho)
- chg_tools.cpp: 1 auto -> std::function<void(int,int)> (kernel)
- chg_symm_detail.cpp: 1 auto -> std::function (build_wspin)

Added #include <functional> to all four files.

* module_charge: wrap lines over 120 chars (Phase 1c)

Break 21 lines exceeding the 120-char limit across 7 files to
eliminate all line_too_long quality-score deductions.

- charge.cpp: 3 WARNING_QUIT/cout lines split
- chg_atomic.cpp: 5 Simpson_Integral/exp/assert lines split
- chg_drho.cpp: 2 conj-product sum lines split
- chg_init.cpp: 1 warning message string split
- chg_mix.cpp: 5 make_twobeta_mix/recip_to_real/if_scf_oscillate lines split
- chg_mix.h: 3 member declaration/comment lines shortened
- chg_symm_detail.cpp: 2 MPI_Recv lines split

* module_charge: remove default parameter from Charge::init_rho (Phase 1d)

Remove the default nullptr values from init_rho's klist and wfcpw
parameters and update the two call sites (esolver_of.cpp,
esolver_double_xc.cpp) that relied on the defaults to pass nullptr
explicitly.

* module_charge: replace raw new/delete with std::vector and unique_ptr (Phase 2a-2d)

Replace all raw new/delete allocations in 4 files with RAII
containers to eliminate raw_new_keyword and unpaired_new_delete
quality-score deductions.

- chg_tools.cpp: 1 new -> std::vector<double> (aux buffer)
- chg_extra.cpp: 4 new -> std::vector<std::vector<double>> (rho_atom
  in extrapolate_charge and find_alpha_and_beta)
- chg_symm_detail.cpp: 14 new -> std::vector (rhog_piece, ig2isz,
  ipsz2ipw, nstnz_start, fftixy2is, rhogtot, ig2isztot, ixyz2ipw
  across reduce_to_fullrhog, rhog_piece_to_all, psymmg, psymmg_soc)
- chg_mix.{h,cpp}: 5 new + 5 unpaired -> std::unique_ptr for
  mixing and mixing_highf members; destructor and init_mixing
  simplified; get_mixing() returns .get()

charge.cpp (18 raw new) deferred to Phase 2e due to wider impact.

* module_charge: replace raw new/delete in Charge with vector-backed storage (Phase 2e)

Replace all 18 raw new and 10 unpaired delete in charge.cpp with
std::vector-backed storage to eliminate raw_new_keyword and
unpaired_new_delete deductions.

- charge.h: add _ptrs_rho, _ptrs_rhog, _ptrs_rho_save, _ptrs_rhog_save,
  _ptrs_kin_r, _ptrs_kin_r_save (std::vector<double*> / complex*),
  and _space_rho_core, _space_rhog_core (std::vector data buffers)
- charge.cpp allocate(): replace new double*[nspin] with vector resize;
  rho = _ptrs_rho.data() preserves double** interface
- charge.cpp init_final_scf(): replace both outer pointer and inner
  data new calls with _space_* vectors
- charge.cpp destroy(): replace delete[] with vector::clear() and
  nullptr assignment

charge.cpp score: 47 -> 69, now passing the 60 threshold.
Module average: 85.0 -> 85.7, 30/33 files passing.

* module_charge: replace std::make_unique with C++11-compatible unique_ptr(new T) (fix)

std::make_unique is a C++14 feature; the repo baseline is C++11.
Replace 4 make_unique calls with std::unique_ptr<T>(new T(...)) to
eliminate the post_cpp11_feature deduction (-40).

chg_mix.cpp score: 0 -> 15, module average: 85.7 -> 86.1.

* module_charge: fix duplicate doc block in charge.cpp init_final_scf

* module_charge: aggregate chgmixing_ks parameters into ScfMixingCtx struct (Phase 3a)

Replace 14-parameter chgmixing_ks with 7-parameter version by
grouping SCF convergence thresholds and status flags into a new
ScfMixingCtx struct, and deriving nrxx from chr.rhopw->nrxx.

- chg_routine.h: define ScfMixingCtx struct (hsolver_error, scf_thr,
  scf_ene_thr, converged_u, drho, oscillate_esolver, conv_esolver)
- chg_routine.cpp: unpack ctx members at function entry
- esolver_ks.cpp: pack ctx before call, unpack after

chg_routine.cpp score: 63 -> 70, too_many_parameters eliminated.

* module_charge: aggregate read_rho_file/read_kin_file parameters into ReadCfg (Phase 3b)

Replace 9-parameter read_rho_file and read_kin_file with 5-parameter
versions by grouping suffix, readin_dir, rank, ofs_running, ofs_warning
into a ReadCfg struct in the anonymous namespace.

chg_init.cpp score: 66 -> 70, too_many_parameters eliminated.

* module_charge: aggregate non_linear_core_correction parameters into NlcCtx (Phase 3c)

Replace 10-parameter non_linear_core_correction with 2-parameter
version by grouping all input data into a new NlcCtx struct.

chg_tools.cpp score: 96 -> 100, too_many_parameters eliminated.

* module_charge: split chg_mix.cpp into init and rho mixing files (Phase 4a)

Move mix_rho_recip, mix_rho_real, and mix_rho (440 lines) from
chg_mix.cpp into a new chg_mix_rho.cpp to eliminate file_too_long
deduction (-10).

- chg_mix.cpp: 727 -> 286 lines (constructor, set_mixing,
  init_mixing, set_rhopw, mix_reset, if_scf_oscillate,
  allocate_mixing_uom, mix_uom)
- chg_mix_rho.cpp: new file, 440 lines (mix_rho_recip,
  mix_rho_real, mix_rho)
- CMakeLists.txt: add chg_mix_rho.cpp to library and test targets

chg_mix.cpp score: 15 -> 60, now passing the 60 threshold.
32/34 files passing, module average improved.

* module_charge: split chg_drho.cpp and decompose inner product functions (Phase 4b)

Move inner_product_recip_rho and inner_product_recip_hartree from
chg_drho.cpp into a new chg_drho_inner.cpp, and decompose each
into per-nspin helper functions to reduce cyclomatic complexity.

- chg_drho.cpp: 520 -> 161 lines (cal_drho, cal_dkin,
  inner_product_real); score 49 -> 97
- chg_drho_inner.cpp: new file, 310 lines; score 100
  - inner_product_recip_rho decomposed into recip_rho_nspin1,
    recip_rho_nspin2, recip_rho_nspin4_mag helpers (CC 29 -> ~5 each)
  - inner_product_recip_hartree decomposed into
    recip_hartree_nspin2, recip_hartree_nspin4_trad,
    recip_hartree_nspin4_angle helpers (CC 37 -> ~5 each)
  - shared coulomb_sum_single extracted
- CMakeLists.txt: add chg_drho_inner.cpp to library and test targets

34/35 files passing, only chg_atomic.cpp remains below 60.

* refactor(module_charge): split atomic_rho and remove ZEROS in charge mixing

chg_atomic.cpp:
- Decompose atomic_rho (CC=60) into per-nspin helpers in
  chg_atomic_inner.cpp; CC reduced to 7, score 40->100.
- Replace all PARAM.inp.nelec/domag/domag_z/test_charge and
  GlobalV::ofs_warning with explicit AtomicRhoCfg parameter.
- Remove unused parameter.h include.
- Add chg_atomic_detail.h declaring detail helpers and RhoG3dCtx.

chg_init/chg_extra/esolver_*:
- Pass AtomicRhoCfg through call sites of atomic_rho,
  extrapolate_charge, and update_delta_rho.

Bug fixes:
- chg_drho_inner.cpp: fix duplicate const (const MixingConfig const&
  -> const MixingConfig&) and add detail:: prefix to helper calls.
- chg_mix_rho.cpp: use mixing.get()/mixing_highf.get() for unique_ptr.
- chg_tools.cpp: fix numeric -> numeric[it] in set_rho_core.

Memory safety / cleanup:
- Replace ModuleBase::GlobalFunc::ZEROS with std::fill in charge.cpp,
  chg_symm_detail.cpp, chg_tools.cpp; remove redundant ZEROS calls
  that precede full overwrites in chg_dmr.cpp and chg_mix_rho.cpp.

* Refactor: remove redundant Charge& overload of cal_rhog_symm_soc

The Charge& overload only forwarded chr.rho/chr.rhog to the raw-array
overload and had a single internal call site. Inline the member access
at that call site and drop the wrapper declaration and definition.

* module_charge: fix stale TITLE/timer labels and drop unused xc_functional.h includes

mix_tau_recip is now a free function in module_charge::detail, so update
its TITLE/timer labels from the legacy "Charge_Mixing" to "module_charge"
to match the convention of other free functions in the directory. Also
remove the unused xc_functional.h includes from chg_tau.cpp and
chg_symm_detail.cpp (label/include cleanup only, no behavior change).

* module_charge: remove redundant #ifdef __MPI guards around parallel wrappers

Parallel_Reduce::reduce_pool and Parallel_Common::bcast_double already
compile to no-op stubs when __MPI is undefined, so the outer guards add
nothing. Remove 11 such guards in chg_tools.cpp, chg_drho.cpp,
chg_drho_inner.cpp, chg_atomic_inner.cpp and chg_mix.cpp.

Guards enclosing raw MPI calls or MPI/serial dual paths are kept
(chg_parallel, chg_symm_detail, chg_routine BP_WORLD bcast, chg_extra.h).

* module_charge: decouple chg_routine from spin_constrain singleton

- forward-declare Plus_U_Base in chg_routine.h instead of including dftu_base.h
- query DeltaSpin mag_converged in ESolver_KS_PW and pass it to chgmixing_ks_pw

* module_charge: remove PARAM dependencies via explicit configuration structs

Remove the last four direct includes of parameter.h in module_charge
(chg_mix, chg_parallel, charge, chg_init) and the implicit PARAM.globalv.ks_run
read in chg_routine. INPUT values are now passed explicitly:

- MixingConfig gains scf_nmax for the drho oscillation history
- reduce_diff_pools/rho_mpi/kin_r_mpi take kpar, all_ks_run, bndpar, nspin,
  out_elf from callers instead of GlobalV::KPAR/PARAM
- Charge::kin_density/allocate/check_rho/renormalize_rho/init_final_scf take
  out_elf/test_charge/nelec as arguments with validation asserts
- new InitRhoCfg aggregates INPUT values for init_rho
- ScfMixingCtx gains ks_run; dm2rho takes nelec and drops its default
  skip_normalize argument per governance rule 5

No behavior change: save_rho_before_sum_band now uses the member nspin
set by allocate, identical to the previously read PARAM.inp.nspin.

* module_charge: restore #ifdef __MPI guards around parallel wrapper calls

The guards removed in 7a0013848 are load-bearing for serial-built unit
tests: source_estate/test strips __MPI from test translation units via
abacus_disable_feature_definitions, but links libbase built with __MPI,
whose explicit Parallel_Reduce instantiations contain real MPI calls.
Unguarded calls in the test TUs therefore bound to MPI_Allreduce and
abort with "called before MPI_INIT", failing MODULE_ESTATE_charge_test
and MODULE_ESTATE_charge_mixing.

Restore all 11 call-site guards in chg_tools.cpp, chg_atomic_inner.cpp,
chg_drho.cpp, chg_drho_inner.cpp and chg_mix.cpp. No behavior change for
MPI or serial production builds.

* Remove dead PAW compensation charge members

nhat, nhat_save in Charge and nhat_mdata in Charge_Mixing have had
no references since #6225 removed the PAW code; drop the orphaned
declarations and update the related comment.

* Refactor: remove unused Charge::prenspin member

prenspin recorded the spin-channel count read from legacy cube charge
files and drove collinear-to-noncollinear rearrangement in init_rho.
After read_rho was replaced by binary read_rhog (#5323, #5362) the
value is neither written nor read anywhere, so drop the dead member.

* Refactor: move Charge::cal_rho2ne/check_rho to module_charge free functions

- Add module_charge::check_rho in chg_tools.{h,cpp} with grid/geometry
  parameters passed explicitly; preserve all branches, thresholds and
  warning/abort messages of Charge::check_rho
- Remove the Charge::cal_rho2ne forwarding wrapper and Charge::check_rho
- Update the three esolver call sites (ks/of/double_xc) to pass rho,
  nspin, rhopw grid sizes and ucell.omega explicitly
- Drop the check_rho stubs in elecstate_pw/base tests and switch
  charge_test to the free cal_rho2ne
- Add test_chg_tools.cpp covering cal_rho2ne, total/spin-polarized
  checks, mismatch warning path and negative-channel aborts

* Refactor: remove redundant Charge::omega_ pointer

- Charge::sum_rho() now reads the cell volume from rhopw->omega, which
  is computed from the same lat0/latvec as ucell.omega and is already
  dereferenced on the same line for nxyz; this also makes the volume
  consistent with the grid rho lives on
- Drop the Charge::omega_ member, its set_omega() setter and the
  chg_init.cpp call site, removing a raw-pointer dependency on the
  UnitCell lifetime; update charge_test accordingly

Verified: MODULE_ESTATE_charge_test and MODULE_ESTATE_chg_tools pass,
elecstate library rebuilds cleanly.

* Remove dead Charge::init_final_scf and allocate_rho_final_scf

init_final_scf has had no production callers since the nscf refactor
(c6ae01236); its only remaining caller was the unit test added in
ba8b7ce9a. After the vector-backed storage refactor it was also a
broken duplicate of Charge::allocate: it never set nspin/nrxx/nxyz/
ngmc and skipped the kin_r buffers. Remove the function, its one-shot
guard flag, and the corresponding test case; destroy() now keys solely
on allocate_rho since vector storage self-manages cleanup.

* Refactor: pass rhopw explicitly to chg_init/chg_routine/chg_extra/chg_symm

Remove implicit reads of chr.rhopw/chr.ngmc from four module_charge files:
- chg_symm.cpp: size kin_g by the rho_basis used for its FFTs
- chg_routine: chgmixing_ks takes const PW_Basis&
- chg_init: orchestrator and four stage helpers take const PW_Basis&;
  the Charge::init_rho member signature is unchanged
- chg_extra: extrapolate_charge/update_delta_rho take const PW_Basis&

Call sites pass *chr.rhopw at the KS boundary or *pw_rhod where the
binding (esolver_fp.cpp chr.set_rhopw(pw_rhod)) makes them identical.
Verified: affected TUs compile and MODULE_ESTATE_charge_extra passes.

* Comments: add TODOs for LCAO+USPP double-grid follow-ups

Record the smooth/dense grid split to revisit if LCAO is ever allowed
with USPP: symmetrize_rho callers pass different grids, and the
ndx/ndy/ndz input path lacks the LCAO guard the ecutrho path has.

* Refactor: replace sticky Charge::cal_elf flag with explicit symm_kin argument

cal_elf was set to true once during ELF output and never reset, so every
later density symmetrization in the same run redundantly symmetrized
kin_r. Replace the mutable workflow flag with an explicit bool parameter
on the Charge& overload of module_charge::cal_rhog_symm:
- ctrl_output_fp passes true right before write_elf consumes kin_r
- symmetrize_rho wrapper and other callers pass XC_Functional::get_ked_flag()

Verified: full incremental build, read_wf2rho unit tests (serial/4 MPI),
write_elf logic test, and tests/01_PW/scf_out_elf (E difference 5e-10 eV,
ELF cube passes CompareFile.py at 3-decimal tolerance).

* Refactor: resolve mixing_tau at config assembly, drop XC dependency from chg_mix

esolver_ks now resolves mix_cfg.mixing_tau = inp.mixing_tau &&
XC_Functional::get_ked_flag() at the single production config assembly
point, so chg_mix/chg_mix_rho no longer query the XC global inside tau
mixing branches (6 sites). test_chg_mix mirrors the resolution in
make_cfg() and sets ked_flag before set_mixing where tau mixing is
expected. Also drop an unused xc_functional.h include from
chg_drho_inner.cpp.

Verified: full incremental build clean; MODULE_ESTATE_charge_mixing
11/11 tests pass; MODULE_ESTATE_charge/chg test suites all pass
(serial + 4-rank MPI).

* Fix: restore complete types in chg_drho_inner.cpp after include removal

Removing xc_functional.h in 87b818f4c broke compilation: the include was
load-bearing transitively, supplying the complete ModulePW::PW_Basis type
and ModuleBase::TITLE. Add the direct includes instead (pw_basis.h,
global_function.h) per IWYU.

Verified: make -j16 exits 0 with full log retained (previous verification
was invalid: a tail pipe masked both the exit code and the errors).

* Refactor: derive tau symmetrization/reduction from kin_r buffer existence

The Charge& cal_rhog_symm overload and rho_mpi/kin_r_mpi queried
XC_Functional::get_ked_flag() (plus a caller-supplied out_elf/symm_kin
flag) to decide whether to touch kin_r. Since Charge::allocate allocates
kin_r exactly when meta-GGA or ELF output needs it, both now check
chr.kin_r != nullptr directly, dropping the XC dependency and the extra
boolean parameters:
- rho_mpi/kin_r_mpi lose the out_elf parameter (2 production, 3 test
  call sites updated)
- the Charge& cal_rhog_symm overload loses the symm_kin parameter
  (ctrl_output_fp, setup_pot, read_wf2rho, update_state_rdmft revert to
  4 arguments); the raw-pointer overload now checks kin_r != nullptr
  only
- module_charge keeps XC references only in charge.cpp, chg_init.cpp,
  chg_drho.cpp (semantic "is meta-GGA" sites, resolved next)

Verified: make -j16 exit 0; 14/14 ctest charge/elecstate/read_wf2rho
tests (serial + 4-rank MPI); tests/01_PW/scf_out_elf integration case
reproduces the reference energy (-194.623411265 eV, diff 5e-10) and the
ELF cube passes CompareFile.py at 3-decimal tolerance.

* Refactor: remove module_xc dependency from module_charge (meta_gga state)

module_charge queried XC_Functional::get_ked_flag() at 5 semantic
"is meta-GGA" sites (tau TF init, tau file read, tau save, tau residual,
tau mixing resolution). Resolve the flag at upper layers instead:
- Charge::allocate takes an explicit meta_gga argument and stores it as
  object state; save_rho_before_sum_band and cal_dkin read it
- InitRhoCfg gains a meta_gga field, filled at the 3 esolver config
  assembly points (ks/of/double_xc)
- delete Charge::kin_density(); 6 esolver call sites inline
  get_ked_flag() || (out_elf[0] > 0) for buffer allocation and pass
  get_ked_flag() as meta_gga; non-SCF allocations pass false
- charge_test mirrors the inline expression

module_charge now has zero references to module_xc.

Verified: make -j16 exit 0 (full log); 14/14 charge/elecstate/
read_wf2rho ctests (serial + 4-rank MPI), including the mGGA tau mixing
and tau-save branches; tests/01_PW/scf_out_elf reproduces reference
energy (-194.623411265 eV, diff 5e-10) and the ELF cube passes
CompareFile.py at 3-decimal tolerance. A SCAN integration case
(205_PW_SCAN) still requires a libxc-enabled build/CI run.

* Fix: allow null rho buffers on ranks with empty real-space grid partition

pack_rho_mag/unpack_rho_mag in chg_rho_detail.h quit whenever any buffer
pointer is null. A rank may legitimately own zero real-space grid points
(nrxx == 0) when the grid is decomposed across more processes than it has
z-slabs (e.g. a 3x3x3 big-cell grid on 4 processes leaves one rank with
no slab); its zero-sized vectors then return null data() pointers even
though the packing loops perform no access. The unconditional check made
LCAO nspin==2 real-space mixing abort with "pack_rho_mag pointer is null"
on such ranks.

Restrict the null-pointer check to n > 0, matching the convention already
used by Parallel_Grid::reduce (only a null buffer with a non-zero size is
a genuine bug). n < 0 remains a hard error. Regression introduced in
d9685d4eb when the inline packing loops were extracted into these helpers.

* Refactor: move rhog_io into module_charge as chg_rhog_io

Relocate source_estate/rhog_io.{h,cpp} to source_estate/module_charge/
under the module_charge namespace, rename include guard to CHG_RHOG_IO_H,
and update the warning tags emitted at runtime. Update both callers
(chg_init.cpp, esolver_fp.cpp) and build files; adapt test_rhog_io.cpp in
place ahead of its move in a follow-up commit. No behavior change.

* Refactor: create module_charge/test with the rhog io unit test

Move test_rhog_io.cpp into module_charge/test/test_chg_rhog_io.cpp with
its support data charge-density.dat, register the new test subdirectory,
and rename the target to MODULE_CHARGE_rhog_io. Remove the migrated
AddTest block from the legacy source_estate/test/CMakeLists.txt.

* Refactor: move charge and charge-extra unit tests into module_charge/test

Rename charge_test.cpp to test_charge.cpp and charge_extra_test.cpp to
test_chg_extra.cpp per the test naming rule, move prepare_unitcell.h
alongside its only users, and register MODULE_CHARGE_charge /
MODULE_CHARGE_extra in the module_charge test CMakeLists. No test data
moves: prepare_unitcell.h only sets file-name strings at runtime, and
the extra test only writes cube files into ./support/.

* Refactor: move mix, parallel and tools unit tests into module_charge/test

Relocate test_chg_mix.cpp (fixing its relative includes), test_chg_parallel.cpp
and test_chg_tools.cpp into module_charge/test, register MODULE_CHARGE_tools /
MODULE_CHARGE_mix / MODULE_CHARGE_parallel with the 4-process mpirun test,
and drop the migrated blocks from the legacy source_estate/test CMakeLists.

* Refactor: rename module_charge test dir to unittests and wire CI for it

Rename source_estate/module_charge/test to unittests (relative CMake
paths are immune to the move). Sync the referencing points: the
add_subdirectory call, the coverage lcov filter (add '*/unittests/*' so
test sources stay excluded from the report), a dedicated Module_Charge
ctest step in test.yml with MODULE_CHARGE added to the catch-all -E
list to avoid double execution, and unittests/ added to the
code_quality_score.py SKIP_DIRS.

* Fix: pass ucell.omega to Charge::sum_rho/renormalize_rho to fix NPT stress

Root cause: commit 34b441e1c ("Refactor: remove redundant Charge::omega_
pointer") changed Charge::sum_rho() to read the cell volume from
rhopw->omega instead of ucell.omega. In variable-cell calculations (NPT),
pw_rho/pw_rhod are NOT rebuilt on cell change (only pw_wfc is), so
rhopw->omega keeps the initial cell volume while ucell.omega is updated
every MD step. The stale volume made sum_rho() return a wrong electron
count, which made renormalize_rho() scale rho by the wrong factor,
corrupting the stress (deviation ~0.002 in 095_PW_NPT) while the total
energy stayed near-correct (variational, second-order sensitive).

Fix: add an explicit omega parameter to Charge::sum_rho() and
renormalize_rho(); all call sites (init_scf, chg_routine, LCAO dm2rho path
through HSolverLCAO/dmToRho, RDMFT update_charge, OFDFT renormalize_psi)
now pass ucell.omega. This mirrors the existing check_rho(..., ucell.omega)
pattern.

Also mark three other rhopw->omega users with BUG(investigate) comments:
get_local_pp_energy, cal_delta_escf, and Makov-Payne correction. These are
pre-existing and were not changed by the refactor; they may have the same
stale-volume issue in NPT and should be investigated separately.

Bisected to 34b441e1c over the 20260916 module_charge refactor branch.

* Fix: add omega arg to remaining dm2rho call sites

Missed four LCAO_domain::dm2rho call sites in the previous commit:
- lcao_set.cpp init_chg_dm (skip_normalize=true, omega unused)
- esolver_dm2rho.cpp
- esolver_ks_lcao_tddft.cpp weight_dm_rho
- module_dm/init_dm.cpp

All now pass ucell.omega.

* Fix: restore HamiltHSMatrix hs declaration in cal_mw_from_lambda

Accidentally removed the line while editing the comment.

* Refactor: merge Charge_Mixing::set_rhopw into set_mixing

Fold the smooth/dense PW_Basis pointer assignment into
Charge_Mixing::set_mixing so grid injection happens together with the
rest of the mixing configuration, and remove the now-redundant
set_rhopw setter. Update the esolver_ks call site and unit tests
accordingly.

* Disable ref_cell_factor != 1.0 and skip 095_PW_NPT test

The reference-cell mechanism (ref_cell_factor > 1) has design problems:
when ref_cell_factor != 1, PW_Basis::lat0/tpiba/G/GGT/omega hold
reference-cell values, but external code (sum_rho, get_local_pp_energy,
cal_delta_escf, makov_payne, wfc IO, DFPT, OFDFT) reads them as
physical-cell quantities, producing wrong charge/energy integration in
variable-cell (NPT) calculations. The bug manifests as a ~16% overestimate
of rhopw->omega on the first MD step (reference-cell volume instead of
actual-cell volume) and silently wrong stress/energy values.

Since properly fixing this requires refactoring PW_Basis to separate
the reference-cell FFT grid (nx/ny/nz) from the physical-cell lattice
quantities (lat0/tpiba/G/GGT/omega), temporarily disable the feature:
- read_input_item_md.cpp: WARNING_QUIT if ref_cell_factor != 1.0
- input_parameter.h: FIXME comment explaining the disable and the
  refactor required to re-enable
- setup_pwrho.cpp / setup_pwwfc.cpp: NOTE comments at the five
  initgrids(ref_cell_factor * ucell.lat0, ...) call sites explaining
  the staleness issue and the initgrids_ref/initgrids_actual split
  needed when the feature is restored
- tests/01_PW/CASES_CPU.txt and CASES_GPU.txt: skip 095_PW_NPT
  (its INPUT sets ref_cell_factor=1.05, which is now blocked)

The 3 BUG(investigate) markers in estate_e_terms.cpp,
elecstate_energy.cpp, and makov_payne.cpp are left in place as
remainders that those call sites also need review when the reference-
cell mechanism is re-enabled.

* module_dm: remove duplicate doc block comments (Phase 1a)

Delete 29 repeated comment lines across 4 files. No logic changes.
- density_matrix.cpp: 17 dups in cal_DMR_td, cal_DMR_full, gamma-only cal_DMR
- density_matrix.h: 4 dups in constructors, cal_DMR_td/full docs, _DMR_grid docs
- density_matrix_io.cpp: 4 dups across init_DMR overloads
- cal_dm_psi.cpp: 4 dups in complex overload

* module_dm: replace auto with explicit types (Phase 1b)

Replace 8 auto occurrences with explicit C++11 types in 3 non-test files.
- density_matrix.cpp: auto& it -> hamilt::HContainer<TR>*& it
- density_matrix_io.cpp: 4x auto& it, 2x auto tau1 -> Vector3<double>
- cal_edm_tddft.cpp: auto Sinv_dev -> ct::Tensor Sinv_dev

* module_dm: unindent preprocessor directives (Phase 1c)

Move 20 indented #ifdef/#pragma/#endif directives to column 0 in
density_matrix.cpp. No logic changes.

* module_dm: wrap lines over 120 chars (Phase 1d)

Break 23 long lines across 4 files. No logic changes.
- density_matrix.cpp: 13 lines (constructors, template specializations,
  func_xyz_to_updown body)
- density_matrix.h: 6 lines (declarations, friend declarations)
- cal_dm_psi.h: 2 lines (psiMulPsiMpi/psiMulPsi declarations)
- cal_dm_psi.cpp: 2 lines (psiMulPsiMpi/psiMulPsi definitions)

* module_dm: convert tab indentation to spaces (Phase 1e)

Convert 39 tab-indented lines to 4-space indentation in 3 files.
- init_dm.cpp: 25 lines
- density_matrix.h: 12 lines
- init_dm.h: 2 lines
No logic changes.

* module_dm: remove default parameters (Phase 1f)

Remove 4 default parameter values from 2 header files and update 35
call sites across 19 files to pass explicit values.
- density_matrix.h: cal_DMR, cal_DMR_td, cal_DMR_full (ik_in=-1 removed)
- cal_edm_tddft.h: print_local_matrix (matrix_name="", rank=-1 removed)
Call sites updated: cal_DMR()->cal_DMR(-1), cal_DMR_td(...)->(...,-1),
cal_DMR_full(&x)->(&x,-1)
No logic changes.

* module_dm: replace raw new/delete with unique_ptr in density_matrix_io (Phase 2a)

- Add clear_DMR() private method to DensityMatrix, consolidating 4
  repeated delete-loop patterns
- Destructor now calls clear_DMR() instead of inline delete loop
- 4 init_DMR overloads: replace delete loops with clear_DMR(), wrap
  7 raw new in std::unique_ptr, use .release() when storing into _DMR
- Add #include <memory> to density_matrix_io.cpp
No logic changes.

* module_dm: replace raw new[]/delete[] with std::vector in cal_edm_tddft (Phase 2b)

Replace 7 raw new[] + 7 delete[] with std::vector in cal_edm_tddft.cpp.
- MPI path: 6x new complex<double>[nloc] -> vector + .data() for ScaLAPACK
- Serial path: 1x new complex<double>[lwork] -> vector + .data() for LAPACK
3 reset(new ...) calls in shared_ptr left unchanged (already owned).
No logic changes.

* module_dm: replace raw new[]/delete[] with std::vector for dmr_tmp_ (Phase 2c)

- Change dmr_tmp_ member from TR* to std::vector<TR>
- Remove delete[] from destructor (vector auto-destructs)
- switch_dmr: nullptr checks -> .empty(), new TR[size] -> .resize(size),
  allocate(dmr_tmp_,...) -> allocate(dmr_tmp_.data(),...)
No logic changes.

* module_dm: eliminate PARAM global dependency in density_matrix.cpp (Phase 3a)

Replace 10 PARAM references with local/parameter alternatives:
- 9x PARAM.inp.nspin -> dm._nspin (already available via DensityMatrix ref)
- 1x PARAM.inp.td_stype==2 -> !phase_hybrid.empty() (semantic equivalent)
No logic changes.

* module_dm: eliminate PARAM dependency in init_dm (Phase 3b)

Introduce Init_DM_Config struct to pass esolver_type, td_stype, nspin,
nelec explicitly instead of reading global PARAM.
- init_dm.h: add struct Init_DM_Config, update init_dm signature
- init_dm.cpp: 5x PARAM.inp.* -> cfg.*, update explicit instantiations
- esolver_ks_lcao.cpp: update call site to pass config struct
No logic changes.

* module_dm: eliminate PARAM.globalv.nlocal in cal_edm_tddft (Phase 3c)

Replace 3x PARAM.globalv.nlocal with pv.nrow (equivalent for LCAO
square matrix). Remove now-unused parameter.h include.
module_dm is now fully free of global PARAM/GlobalV/GlobalC deps.
No logic changes.

* module_dm: split density_matrix.cpp into density_matrix + dmr_cal (Phase 4a)

Extract DMR calculation functions (cal_DMR, cal_DMR_td, cal_DMR_full
templates and specializations) from density_matrix.cpp into new
dmr_cal.cpp. Update CMakeLists.txt.
- density_matrix.cpp: 707 -> 290 lines
- dmr_cal.cpp: new, 430 lines
No logic changes.

* module_dm: split cal_edm_tddft.cpp into cal_edm_tddft + cal_edm_tddft_lapack (Phase 4b)

Extract cal_edm_tddft_tensor and cal_edm_tddft_tensor_lapack (plus the
explicit template instantiations for CPU/GPU) into a new translation unit
cal_edm_tddft_lapack.cpp so cal_edm_tddft.cpp keeps only print_local_matrix
and the ScaLAPACK-based cal_edm_tddft driver. cal_edm_tddft.cpp drops from
819 to 309 lines. Wire the new file through source_estate/CMakeLists.txt.

* module_dm: cleanup after Phase 4 split (dead code, typo, build fixes)

- Remove unused print_local_matrix and cal_edm_tddft_tensor from
  cal_edm_tddft.{h,cpp} and cal_edm_tddft_lapack.cpp (no callers).
- Fix missing #endif for the __CUDA guard in cal_edm_tddft_lapack.cpp
  that previously swallowed the namespace close.
- Fix Tk -> TK typo in dmr_cal.cpp (2 occurrences).
- Update test_cal_dm_r.cpp to pass explicit -1 to cal_DMR / cal_DMR_full
  after Phase 1f removed the default parameters.
- Add dmr_cal.cpp to test CMakeLists for MODULE_ESTATE_dm_cal_DMR_test,
  dftu_lcao_test, and deepks_unit_support so the cal_DMR specialization
  is linked (Phase 4a moved it out of density_matrix.cpp).

* Fix: close_kerker_gg0 actually disables Kerker; drop dead mixing_gg0 members

The chg_precond refactor (commit 6d127d517) made the Kerker kernels read
cfg_ (immutable INPUT snapshot) instead of Charge_Mixing members, but
close_kerker_gg0() kept writing the now-dead mixing_gg0/mixing_gg0_mag
members. As a result, the non-separate-loop EXX path in exx_lri_interface.hpp
silently failed to disable Kerker after convergence.

Fix: add a kerker_disabled_ flag on Charge_Mixing that the mix_rho_recip/
mix_rho_real screening lambdas short-circuit on. The flag lives on the
object, not in cfg_, so the immutable INPUT snapshot invariant is preserved.

Also drop the now-dead members mixing_gg0/mixing_gg0_mag/mixing_gg0_min/
mixing_angle/mixing_dmr and the get_mixing_gg0() getter; set_mixing/init_mixing
now read these from cfg_ directly. Add CloseKerkerGg0DisablesScreenReal
regression test that compares close_kerker_gg0() output against the
cfg.mixing_gg0=0 baseline and proves the flag is load-bearing.

* Fix: relax over-strict null-buffer asserts for empty grid partitions

reduce_diff_pools and Parallel_Grid::reduce_across_pools still forbade
null buffers unconditionally, contradicting the rule documented at
parallel_grid.cpp:355-360. A rank with nrxx == 0 may legitimately hold
a null rho/kin_r pointer; the MPI calls below use count 0 and ignore
the buffer. Align both call sites with the documented rule.

* Fix: relax over-strict null-buffer assert in ParaRgridWorld::reduce_across_pools

Same pattern as the previous fix: a rank with nrxx == 0 legitimately
holds a null buffer, and MPI_Allreduce with count 0 ignores it.
Align with the rule documented at parallel_grid.cpp:355-360.

* Fix: allow nnr == 0 in DMR mixing for empty MPI partitions

nnr is local to each MPI rank and may legitimately be zero when no
atom pairs survive the cutoff on that rank. The previous check
aborted DMR mixing for such distributions, whereas the historical
implementation allowed empty blocks. Relax the guard in
check_dmr_inputs() and init_mixing_dmr() to reject only negative
nnr, and require non-null DMR buffers only when nnr > 0, matching
the established nrxx == 0 convention in module_charge.

* Fix: split reciprocal rho copy from real-space |m| rescale in mix_rho_recip

The nspin==4 && mixing_angle>0 branch of mix_rho_recip mixed two
distinct operations in one loop bounded by npw, but rho_magabs is
sized nrxx (real-space) and the new |m| is written back by
recip2real into rho_magabs[0..nrxx-1]. Reading rho_magabs[npw+ig]
goes out of bounds once npw+ig >= nrxx (AddressSanitizer reproduces
with nrxx=125, npw=93) and the loop bound npw leaves the real-space
tail [npw, nrxx) of {mx,my,mz} unscaled. Split into two loops: the
reciprocal rho copy stays bounded by npw, the magnetization rescale
is bounded by nrxx and reads rho_magabs[ir].

* Refactor: remove unused Charge_Mixing::conserve_setting

conserve_setting() was introduced by 420f1ad00 (DeltaSpin feature
merge, 2026-06-15) but never wired up: no production caller, no
test reference, and the DeltaSpin module does not touch
Charge_Mixing. Drop the dead declaration per the project rule that
unused functions and their tests be removed.

* Refactor: drop dead Charge_Mixing::tpiba2 member

tpiba2 was declared in chg_mix.h but never assigned by set_mixing()
nor read anywhere in the module. Grep across the whole source tree
confirms all tpiba2 references are either ucell.tpiba2 (a separate
UnitCell member) or local variables in unrelated modules. The
Charge_Mixing class never computed or used its own tpiba2 pointer;
only tpiba is consumed by the stateless Kerker kernels via
mix_rho_recip/mix_rho_real. Remove the dead declaration.

* Refactor: route Charge_Mixing getters through cfg_

get_mixing_mode(), get_mixing_beta(), get_mixing_ndim() previously
returned the legacy mirror members that set_mixing() kept in sync
with cfg_ by hand. With cfg_ now treated as the immutable INPUT
snapshot, route the public getters through cfg_ directly so there
is a single source of truth for INPUT parameters. External callers
(esolver_ks_lcao, lcao_others, pw_others) are unaffected since
signatures are unchanged. The legacy members remain in place for
now; they are dropped in a later step after internal readers are
migrated.

* Refactor: init_mixing constructs Mixing from cfg_ not legacy mirrors

init_mixing() branched on this->mixing_mode and passed
this->mixing_ndim/mixing_beta to the Broyden/Pulay/Plain_Mixing
constructors. These legacy mirrors were kept in sync with cfg_
manually by set_mixing(). Route through cfg_ directly so cfg_
remains the single source of INPUT parameters. The Mixing objects
themselves still copy beta/ndim into their own members at
construction; that is a one-time snapshot and not a continuous
sync surface, so it is left untouched.

* Refactor: mix_rho_recip/mix_rho_real read mixing_beta from cfg_

Both mix_rho_recip and mix_rho_real built the twobeta_mix functor by
reading this->mixing_beta / this->mixing_beta_mag, which are legacy
mirrors that set_mixing() kept in sync with cfg_. Route the six
construction sites through cfg_.mixing_beta / cfg_.mixing_beta_mag
so cfg_ is the single source of INPUT parameters consumed by the
mixing logic. Behavior is unchanged since the mirrors and cfg_
hold identical values after set_mixing().

* Refactor: set_mixing stops mirroring cfg_ into legacy members

set_mixing() copied mixing_mode, mixing_beta, mixing_beta_mag,
mixing_ndim from cfg into legacy mirror members, then validation
and logging read from the mirrors. Now that all internal readers
(init_mixing, mix_rho_recip, mix_rho_real, getters) read from
cfg_, the mirror writes are dead work. Drop them and route
validation and log output through cfg_ directly. omega and tpiba
remain pointer members because they alias external runtime state
(cell volume, lattice constant) that changes across SCF iterations
and so do not belong in MixingConfig (an immutable INPUT snapshot).

* Refactor: drop legacy Charge_Mixing mirror members; cfg_ is single source

Drop mixing_mode, mixing_beta, mixing_beta_mag, mixing_ndim mirror
members. After the previous commits every internal reader (getters,
init_mixing, mix_rho_recip, mix_rho_real, set_mixing validation
and log output) routes through cfg_, so the mirrors are dead state
that set_mixing() no longer writes. cfg_ is now the single source
of truth for INPUT mixing parameters.

Update test_chg_mix.cpp accordingly: the two assertions that
reached directly into CMtest.mixing_beta_mag and CMtest.mixing_mode
now read CMtest.get_mixing_config().mixing_beta_mag and
CMtest.get_mixing_mode(), matching the public API used by the
other assertions in the same block. No production caller accessed
these members directly (esolver_ks_lcao, lcao_others, pw_others
all used the getters), so the change is test-only on the consumer
side.

* Refactor: drop NSDMI from MixingConfig to force explicit construction

The non-static data member initializers in MixingConfig provided
plausible-looking defaults (e.g. mixing_beta=0.8, mixing_mode=
"broyden") that silently masked forgotten fields when a new field
was added but not wired up at construction sites. With the
defaults removed, every construction site must use aggregate
initialization (or copy-assign from a fully-initialized instance),
and a missing field yields value-initialized (zero/empty) members
that are far more likely to trip a test than the old defaults.
Combined with -Wmissing-field-initializers promoted to error in
the next commits, adding a field to MixingConfig without updating
all aggregate-initialization sites becomes a compile error.

* Refactor: aggregate-init MixingConfig in esolver_ks with pragma guard

Convert the 17-line field-by-field assignment of mix_cfg into a
single aggregate initialization in declaration order. Wrap it in
#pragma GCC diagnostic error "-Wmissing-field-initializers" so
that adding a field to MixingConfig without updating this list
becomes a compile error rather than silently using a default.
Each initializer is annotated with the field name it corresponds
to, making the declaration-order dependency auditable at a glance.

* Refactor: aggregate-init MixingConfig in test_chg_mix with pragma guard

Convert make_cfg()'s 17-line field-by-field assignment into a
single aggregate initialization in declaration order, matching
the esolver-side change. Wrap in the same
#pragma GCC diagnostic error "-Wmissing-field-initializers" so
that adding a field to MixingConfig without updating the test
helper is also a compile error. Both construction sites (esolver
and test) now fail at compile time if a field is missing, closing
the maintenance gap where a new field could silently fall back to
a default value.

* Fix: fail-fast guards in Charge_Mixing and update chg_mix tests

Add validation to turn latent misuse (skipped set_rhopw/set_mixing)
into clear WARNING_QUIT errors instead of null dereference or heap
corruption:
- init_mixing rejects a null rhopw
- if_scf_oscillate checks scf_nmax > 0 and iteration range
- mix_rho validates chr/chr->rhopw and the grid pointers

Fix three chg_mix unit tests that read cfg_ before set_mixing, which
caused a SIGSEGV in SCFOscillationTest and assertion failures in the
two inner-product tests.

* test(module_charge): add unit tests for chg_uspp and chg_dmr

Add test_chg_uspp.cpp covering split_dgrid/merge_dgrid (normal split,
round-trip, nspin=1/2, empty high-frequency/smooth boundaries, and
input-validation abort paths).

Add test_chg_dmr.cpp covering init_mixing_dmr/mix_dmr (nspin=1/2/4
mixing with Plain_Mixing analytically verified, empty-partition null
buffer allowance, and input-validation abort paths).

Wire both targets into unittests/CMakeLists.txt.

* test(module_charge): add unit tests for chg_precond, chg_drho, chg_drho_inner, chg_mix_rho

- test_chg_precond.cpp: kerker_screen_recip/real (early return, nspin=1/2/4
  filter, nspin=4 with mixing_angle resize, real-space matches reciprocal).
- test_chg_drho.cpp: inner_product_real, cal_drho real-space path
  (nspin=1/2/4+domag_z), cal_dkin (meta_gga false/true).
- test_chg_drho_inner.cpp: inner_product_recip_rho and
  inner_product_recip_hartree for nspin=1 with a single G component,
  analytically verified against the Coulomb metric.
- test_chg_mix_rho.cpp: mix_rho abort paths (null chr/chr->rhopw, unset
  rhopw, double_grid without rhodpw) and real-space plain mixing value.

Wire all four targets into unittests/CMakeLists.txt.

* test(module_charge): add unit tests for chg_symm, chg_symm_detail, chg_atomic, chg_atomic_inner

- test_chg_symm.cpp: symmetrize_rho / cal_rhog_symm / cal_rhog_symm_soc
  no-op paths when symm_flag == 0, for nspin=1 and nspin=4.
- test_chg_symm_detail.cpp: psymmg and psymmg_soc idempotence on a
  manually built D_4 point group over a serial cubic PW_Basis.
- test_chg_atomic_inner.cpp: compute_rhoatm USPP direct-copy branch and
  NCPP integrate+scale-to-zv branch (Gaussian rho_at with known analytic
  integral); normalize_and_check renormalizes uniform density to nelec.
- test_chg_atomic.cpp: atomic_rho ntype==0 path (skips atom loop) and
  spin_number_need==3 abort path.

Wire all four targets into unittests/CMakeLists.txt.

* test(module_charge): add chg_tau/chg_routine/chg_init tests; drop spurious XC_Functional stubs

Fourth batch of module_charge unit tests:
- test_chg_tau.cpp: mix_tau_recip abort paths (null chr/grid/mixing, nspin<1,
  double_grid without high-f mixer) and non-double-grid plain mixing value.
- test_chg_routine.cpp: chgmixing_ks_pw/lcao iter==1 restart-step setup, and
  chgmixing_ks convergence branches (conv_esolver true / drho<hsolver_error
  skip mix_rho).
- test_chg_init.cpp: init_rho "wfc" with null wfcpw abort, and "atomic" with
  ntype==0 + meta_gga Thomas-Fermi tau initialization.
Wire all three targets into unittests/CMakeLists.txt.

Cleanup: remove the XC_Functional::func_type /…
Growl1234 and others added 19 commits September 25, 2026 04:50
* Docs: Fix Sphinx build warnings

* Improve comments and fix program_id formatting

Updated comments for clarity and corrected program_id format.

---------

Co-authored-by: Levi Zhou <31941107+ZhouXY-PKU@users.noreply.github.com>
Fix memory leak where pdsxk and tmp3 were allocated but never deleted.
Replace all raw complex<double>* work buffers with std::vector to
ensure automatic cleanup on all exit paths.

Fixes #7554

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
* Fix spin-channel mismatch in LR complex transition dipole (length gauge)

LR_Spectrum<complex<double>>::cal_transition_dipole_istate_length looped
over all nspin_x spin channels of the transition density matrix but
called get_DMR_real_imag_part() without selecting a channel, so the
function copied every channel of DM_trans (size nspin_x) into the
single-channel DM_trans_real_imag. For nspin_x==2 (open-shell/spin-
polarized, multi-k) this hits an assertion failure (or out-of-bounds
access in release builds), and even when it doesn't crash it double-
counts the merged density on each iteration of the outer loop.

Add an overload of get_DMR_real_imag_part() that copies a single named
spin channel into the (always single-channel) DMR_real, and use it from
the two call sites in cal_transition_dipole_istate_length so each outer
loop iteration only processes its own channel, matching the double
(gamma-only) specialization's behavior.

Verified: reproduced the original assertion failure with gdb on the
open-shell (nspin=2) multi-k length-gauge path, confirmed the fix
resolves it, and confirmed the fixed multi-k oscillator strength /
transition dipoles match the gamma-only result exactly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Fix segfault in LR real/imag HContainer helpers under MPI-parallel runs

get_DMR_real_imag_part() (both overloads) and set_HR_real_imag_part()
dereferenced the result of HContainer::find_pair(ia, ja) without a
null check. Under real multi-rank MPI parallelism, the HContainer is
distributed 2D block-cyclic, so find_pair() legitimately returns
nullptr for atom pairs not owned by the current rank — dereferencing
that pointer segfaults.

Reproduced with gdb on tests/08_EXX/54_GO_ULR_HF (gamma_only=0, KPT
1 1 1, mpirun -np 2): SIGSEGV in get_DMR_real_imag_part, called from
OperatorLRHxc::grid_calculation during the Casida eigenvalue solve —
a different call site/crash than the spin-channel-mismatch bug fixed
in the previous commit, and only reproducible with more than one MPI
rank (a single-rank run owns every atom pair locally, masking the
bug).

Skip atom pairs not present on the local rank, matching the existing
pattern used elsewhere in the codebase for MPI-parallel HContainer
access. Verified: the same 54_GO_ULR_HF case now runs cleanly under
mpirun -np 2 (gdb shows a clean exit, no crash) and its excitation
energies exactly match both the gamma_only=1 run and result.ref.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Adapt new get_DMR_real_imag_part overload to module_dm refactor

Upstream #8000 renamed elecstate::DensityMatrix to module_dm::DensityMatrix
and get_DMR_vector/get_DMR_pointer to get_dmr_vec/get_dmr_ptr; update the
spin-index overload added in this branch accordingly.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* Refactor: scaffold module_proj for onsite projector extraction

Create an empty source_pw/module_proj OBJECT library and wire it into the
build (add_subdirectory before module_pwdft; final link after module_pwdft).
A placeholder .cpp keeps the target valid until the first real source moves
in. No functional change; full CPU build and governance check pass.

* Refactor: move radial_proj into module_proj

Migrate RadialProjection::RadialProjector (SBT radial projector) from
module_pwdft to the new module_proj. Pure file relocation: update include
paths in onsite_proj.h and radial_proj_test.cpp, rewire both CMakeLists and
the Makefile.Objects VPATH. Remove the placeholder source.

No behavior or interface change, so no docs/parameters update is required.
Verified: full CPU build links, MODULE_PW_radial_proj_test passes.

* Refactor: move onsite_proj_print into module_proj

Migrate the onsite projector print helpers (print_orb_chg, print_mag_table,
print_proj_status) from module_pwdft to module_proj. Pure relocation: update
the three include paths and both CMakeLists. Makefile.Objects needs no change
since module_proj is already on the VPATH.

No unit test exists for these print helpers; no behavior or interface change,
so no docs update is required. Verified: full CPU build links.

* Refactor: move Onsite_Proj_tools into module_proj

Migrate the onsite projector computation backend (cal_becp / cal_dbecp /
cal_force / cal_stress, six files) from module_pwdft to module_proj. Update
the self/nonlocal_maths includes to full paths and rewire both CMakeLists.

Onsite_Proj_tools still includes module_pwdft headers (vnl_pw.h, kernels/
{force,stress}_op.h); this temporary reverse dependency is tolerated and
recorded in module_proj/CMakeLists.txt. kernels/onsite_op stays in pwdft
since it is used by op_pw_proj.cpp, not by the tools. nonlocal_maths.hpp
also stays in pwdft (shared by seven files). OBJECT libraries link fine.

No behavior or interface change; verified by full CPU build linking.

* Refactor: move OnsiteProjector core into module_proj

Migrate the OnsiteProjector class (onsite_proj.h/.cpp, onsite_proj_init.cpp,
onsite_proj_overlap.cpp) from module_pwdft to module_proj, and update the nine
include sites across lcao/io/pwdft. The DFT+U/DeltaSpin force & stress member
functions (onsite_proj_force_stress.cpp) stay in module_pwdft since they
depend on Plus_U_Base; the class declaration now lives in module_proj while
those definitions remain in pwdft, linked via the OBJECT libraries.

The singleton and public API are unchanged. No behavior or interface change,
so no docs update is required. Verified: full CPU build links, and the
deltaspin ctest set (5 tests incl. deltaspin_pw_test) passes.

* Refactor: drop dead commented code in tabulate_atomic

The commented-out STAGE 1/STAGE 2 body of OnsiteProjector::tabulate_atomic is
obsolete: the actual tabulation is done by Onsite_Proj_tools. Remove the dead
block and keep only the k-point dimension bookkeeping, with a comment
recording where the real work happens. No behavior change; full CPU build links.

* Refactor: extract occ_from_proj into source_estate

The per-projector 2x2 occupation block accumulation
rho^{ss'} = sum_i w_i conj(proj^s) proj^{s'}
was previously inlined in OnsiteProjector::cal_occupations and
duplicated in DeltaSpin's accumulate_Mi_from_becp. Extract it as a
free function elecstate::occ_from_proj in source_estate/occ_comput so
both call sites can share it; this commit only adds the function and
its unit tests, call-site migration follows in later commits.

Verification: ctest -R MODULE_ESTATE_occ_comput passes (5 cases).
No INPUT parameter change, docs not required.

* Refactor: migrate cal_occupations to occ_from_proj

Replace the inline occupation accumulation loop in
OnsiteProjector::cal_occupations with the shared free function
elecstate::occ_from_proj extracted in the previous commit. The
behavior is identical: same weights, same spin-channel placement
(isk for nspin=2), same even split for nspin=1. The local variables
proj_p, wg_ik, isk and nat are computed before the call so no
conditional expression appears in the argument list.

Verification: module_proj target builds; numerical equivalence is
covered by the oracle-based unit tests in test_occ_comput.cpp.
Runtime check pending on a full abacus build.
No INPUT parameter change, docs not required.

* Refactor: migrate accumulate_Mi_from_becp to occ_from_proj

Rebuild spinconstrain::accumulate_Mi_from_becp on the shared core
elecstate::occ_from_proj instead of its own inline becp loops. The
public signature (including spin_sign) is unchanged, so call sites in
deltaspin_pw_mi.cpp need no modification. Internally the function now
computes the per-projector 2x2 occupation blocks once, then aggregates
them into per-atom magnetic moments:
  npol=2 (nspin=4): Mi = sum_iprj pauli_to_moment(block_iprj)
  npol=1 (nspin=2): Mz = sum_iprj (occ[0] - occ[3]) == weight*occ*spin_sign
spin_sign is mapped to isk (+1 -> 0, -1 -> 1) for the shared core;
nspin=1 never occurs in DeltaSpin.

Two new unit tests (RealFunction_Npol1/Npol2_MatchesOracle) call the
real function and require agreement with the pre-refactor oracle loops.
deltaspin_core_test now links mi_tools.cpp and occ_comput.cpp.

Verification: ctest -R deltaspin_core_test passes (24 cases, incl. the
2 new ones). No INPUT parameter change, docs not required.

* Docs: add governance rule on named locals in argument lists

Record coding rule 14 in AGENTS.md: do not write conditional or
computed expressions in a function's argument list; assign to a named
local variable first and pass that variable. This was agreed during the
occ_from_proj refactor (commits 34e7c5cb9, 1e381236f) and is applied in
those call sites.

* Fix: add occ_comput.o to Makefile.Objects link list

The Makefile build was missing occ_comput.o, causing undefined
references to elecstate::occ_from_proj from onsite_proj_overlap.o
and mi_tools.o at link time. The CMake build already includes it.

* Refactor: convert RadialProjector helpers to free functions and migrate tests

- Convert _build_backward_map, _build_forward_map, _build_sbt_tab
  (type-wise overload), and _mask_func from RadialProjector static
  members to RadialProjection namespace free functions.
- Remove unused sbfft() declaration and _do_mask_on_radial empty
  implementation along with ~100 lines of commented-out legacy code.
- Slim radial_proj.h by dropping unitcell.h and pw_basis_k.h includes;
  add matrix.h and <tuple> for free function signatures.
- Update onsite_proj_init.cpp call sites to the new free function API.
- Remove dead RadialProjector rp_ member from OnsiteProjector and drop
  the now-unneeded radial_proj.h include from onsite_proj.h.
- Migrate radial_proj unit test from module_pwdft/test to
  module_proj/unittests, renamed to test_radial_proj.cpp per governance
  naming rules, with API calls updated.
- Register the test in module_proj/unittests/CMakeLists.txt under
  BUILD_TESTING and remove the old registration from
  module_pwdft/test/CMakeLists.txt.

Verification: python3 tools/03_code_analysis/agent_governance_check.py
--staged (no findings). Compile and runtime tests not run per instruction.

* Refactor: remove duplicated read_abacus_orb, reuse ModuleIO version

- Remove OnsiteProjector::read_abacus_orb member function and its
  explicit instantiations; the implementation duplicated the existing
  ModuleIO::read_abacus_orb in source_base/module_out/orb_io.h.
- Update init_proj to call ModuleIO::read_abacus_orb directly.
- Drop the now-unneeded parallel_common.h include from
  onsite_proj_init.cpp.

Verification: python3 tools/03_code_analysis/agent_governance_check.py
--staged (no findings). Compile and runtime tests not run per instruction.

* Refactor: remove dead code and debug comments from overlap_proj_psi

- Delete ~40 lines of commented-out legacy gemm implementation in
  overlap_proj_psi.
- Delete debug std::cout comment lines in cal_occupations.

Verification: python3 tools/03_code_analysis/agent_governance_check.py
--staged (no findings). Compile and runtime tests not run.

* refactor(module_proj): fix UB in transfer_gcar by unpacking Vector3 elementwise

The caller passed &(gcar[ik*npwk_max].x) and transfer_gcar copied it via
gcar_tmp.assign(gcar_in, gcar_in + 3*npw_max), which assumes Vector3<double>
is a contiguous 3-double POD. The standard does not guarantee the memory
layout of Vector3, so this was undefined behavior (as noted in the code
comments).

Change transfer_gcar to take const ModuleBase::Vector3<FPTYPE>* and unpack
x/y/z elementwise into the contiguous buffer. This makes the copy well
defined regardless of Vector3's internal layout.

* Fix: add missing includes in module_proj for standalone compilation

- radial_proj.h: include source_base/realarray.h for ModuleBase::realArray
- radial_proj.cpp: include <algorithm> for std::max_element/transform/for_each
- onsite_proj_init.cpp: include radial_proj.h for RadialProjection namespace

These were previously satisfied via transitive includes that were removed
during header dependency cleanup.

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
* Fix BLACS grid leak in Parallel_2D (issue #8003)

Problem: Parallel_2D::init() creates BLACS grids via Csys2blacs_handle()
and Cblacs_gridinit() but never releases them. When kpar > 1 in LCAO
calculations, each SCF iteration leaks BLACS grids, eventually causing
"Too many communicators" error after ~4094 iterations.

Solution: Add ownership semantics to Parallel_2D:
- init() creates grid and sets owns_blacs_ctxt_ = true
- set() borrows existing grid and sets owns_blacs_ctxt_ = false
- Destructor releases grid only if owns_blacs_ctxt_ is true
- Move semantics properly transfer ownership to prevent double-free

Also:
- Remove explicit Cblacs_gridexit() calls in diago_hs_para() since
  destructor now handles cleanup automatically
- Remove Cblacs_exit(1) in esolver_ks_lcao.cpp to prevent
  destroying BLACS environment before Parallel_2D destructors run

* Fix BLACS grid leak in Parallel_2D and related code

This commit fixes the BLACS grid leak issue (#8003) and related
resource management problems:

1. Parallel_2D: Add RAII-style BLACS grid management
   - Add destructor, move constructor, move assignment
   - Add release_blacs_grid() to properly release owned grids
   - init() sets ownership flag, set() borrows without owning

2. Parallel_K2D: Replace raw pointers with std::unique_ptr
   - Automatic cleanup, exception-safe
   - Fix copy-initialization to direct construction

3. module_genelpa/utils: Remove unused initBlacsGrid()

4. ELPA_Solver: Add documentation about BLACS context ownership

5. Test fixes:
   - blacs_connector_test: Add Cblacs_gridexit() calls
   - tddft_test: Use Cblacs_gridexit() instead of Cblacs_exit()
   - single_r_io_test: Add Parallel_2D destructor mock
   - parallel_k2d_test: Fix copy-initialization

Fixes: #8003

* Fix C++11 compatibility: replace std::make_unique with unique_ptr::reset

std::make_unique is a C++14 feature; ABACUS requires C++11.

* Fix blacs_ctxt lost in Parallel_2D move assignment

In operator=(Parallel_2D&&), rhs.blacs_ctxt was reset to -1 before being
copied to this->blacs_ctxt, so any moved Parallel_2D (e.g. via
paraX_.emplace_back(std::move(px)) in esolver_lr_lcao_bse) ended up with
blacs_ctxt == -1. Downstream setup_2d_division then called
Cblacs_gridinfo(-1), leaving coord {-1,-1} and making descinit_ fail with
"DESCINIT parameter number 6 had an illegal value", followed by
std::length_error from vector resize. Copy the context before resetting rhs.

* Fix MODULE_DFTU_folding segfault from Parallel_2D ODR/ABI mismatch

The folding test was built without __MPI while the base library is built
with __MPI. After commit 8505a21 added the __MPI-only member
owns_blacs_ctxt_ and a non-trivial destructor to Parallel_2D, the class
layout now differs between the two TUs (sizeof 136 vs 184). The test TU
used the base library's destructor on an object laid out by its own
(no-__MPI) view, so the four std::vector members were read out of bounds
and free() was called on a wild pointer -> SIGSEGV in TearDown.

Fix by compiling the test with __MPI to match the base library, and by
linking the real parallel_orbitals.cpp instead of stubbing the
Parallel_Orbitals ctor/dtor, so the Parallel_Orbitals/Parallel_2D layout
is consistent across TUs.

Verified: ctest -V -R MODULE_DFTU_folding -> 2/2 passed.

* Fix Parallel_2D::set destroying its own BLACS grid on context reuse

When the blacs_ctxt passed to set() is the context owned by the object
itself (e.g. the block-size fallback pv.set(..., pv.blacs_ctxt) in
lcao_init_basis.cpp, hit by kpar = 1 runs), release_blacs_grid() called
Cblacs_gridexit on the very grid about to be reused, and ownership was
dropped, causing a DESCINIT error and invalid descriptors. Preserve the
grid and its ownership in this case and only rebuild the distribution
info. Add a unit test (SetWithOwnCtxt) covering this path.

* Fix set_serial leaving stale blacs_ctxt on borrowed contexts

release_blacs_grid() is a no-op for borrowers, so set_serial() on an
object using a borrowed context kept the old blacs_ctxt: comm() then
returned a live communicator instead of MPI_COMM_NULL, and the handle
dangled once the owner destroyed the grid. Clear blacs_ctxt
unconditionally when switching to serial mode. Add a unit test
(SetSerialClearsBorrowedCtxt) covering the borrower path.

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
The ASE interface read_stru checked for the LATTICE_PARAMETER block but
then indexed blocks['LATTICE_PARAMETERS'] (extra 'S'), so STRU files
using LATTICE_PARAMETER instead of LATTICE_VECTORS raised a KeyError
(issue #7555). Read from blocks['LATTICE_PARAMETER'][0] instead, since
block values are lists of lines. Add a unit test covering a STRU with
LATTICE_PARAMETER and no LATTICE_VECTORS.

Verified: python3 -m unittest abacuslite.io.generalio.TestAbacusCalculatorIOUtil -v
(10 tests OK, 2 pre-existing skips)

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
* Fix #7556: preserve all atoms when saving STRU with repeated species

Cell._save_stru() rebuilt species_dict[symbol] on every atom, resetting
the species atom list each time. For multiple atoms of the same species
only the last atom survived, and natom was tied to the pp_type branch
instead of counting appended atoms.

Initialize each species entry only once and increment natom for every
appended atom independently of pp_type. Add a regression test that
appends atoms of an existing species and checks the saved STRU reloads
with the full atom list.

* Fix #7558: call RadialCollection methods correctly in overlap_generator

The PyABACUS wrapper exposes rcut_max and lmax as methods, not
properties, and there is no lmax_ method. The old code treated them as
attributes, so overlap generation failed before producing matrices.

Use self.orb.rcut_max(), self.orb.lmax(), and self.orb.lmax(it).

* delete overlap_generator.py

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
…ite (#8025)

* fix: guard Binstream write path against unopened files and failed fwrite

Fixes #7562.

The binary wavefunction writers (wfc_nao_write2file and
wfc_nao_write2file_complex) used to only print a warning when the output
file could not be opened, then kept writing through a null FILE pointer,
which could crash or produce corrupt files.

Binstream::operator<< and Binstream::write now check that the stream is
open and verify the fwrite return value, matching the existing read-path
behavior. The call sites in write_wfc_nao.cpp now terminate with
WARNING_QUIT on open failure instead of continuing.

Tests added:
- BinstreamTest.WriteToUnopenedFile
- ModuleIOTest.WriteWfcNaoBinaryOpenFail
- ModuleIOTest.WriteWfcNaoComplexBinaryOpenFail

* fix: detect delayed write failures in Binstream and route all errors through WARNING_QUIT

fwrite() reports success while data is still buffered, and both close()
and the destructor used to ignore fclose() failures, so a delayed write
error (e.g. RLIMIT_FSIZE, full disk) could be silently dropped: with
RLIMIT_FSIZE=1 the NAO binary writers returned normally with only one
byte on disk.

Changes:
- operator<< and write() now fflush() after fwrite() so buffered write
  errors surface at the call site; the array overload had the same gap.
- close() checks the fclose() result and WARNING_QUITs on failure, and
  is now a safe no-op on an unopened stream (fclose(NULL) was UB).
- ~Binstream() checks fclose() but only WARNINGs, since a destructor
  must not terminate the program.
- All error paths (read/write, scalar/array) now exit via
  ModuleBase::WARNING_QUIT (exit code 1, warning.log, MPI-safe) instead
  of std::cout + exit(0/1).
- operator>> and read() now reject an unopened stream instead of
  calling fread(NULL), matching the write path.
- open() closes any previously opened file first instead of leaking the
  old handle; copy/assignment are deleted (FILE* ownership would
  double-close).
- Error messages: fix "didn't be" grammar and the misleading "dynamic
  memory" wording for array overloads.

Tests added/updated:
- BinstreamTest.DelayedWriteFailureDetected: RLIMIT_FSIZE=1 + ignored
  SIGXFSZ, asserts write() exits with code 1.
- BinstreamTest.ReadFromUnopenedFile, BinstreamTest.CloseFailureDetected.
- Existing death tests updated to ExitedWithCode(1) and the "!NOTICE!"
  marker.

No docs update needed: no INPUT parameter or user-facing interface
behavior changed, only failure handling of binary I/O.

Verified: make MODULE_BASE_binstream && ctest -R MODULE_BASE_binstream
(6/6 passed); full make -j8 with no errors.

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
…n addition #6458 is fixed (#8023)

* Fix: validate cube file parsing before copying charge data (#7563)

read_vdata_palgrid() ignored the bool return of read_cube() and used
the parsed dimensions and data unconditionally, so a malformed or
truncated cube file could lead to out-of-bounds access in memcpy or
trilinear_interpolate().

- read_cube() now checks the stream state after each parsing stage,
  rejects negative natom and non-positive grid dimensions, guards the
  nx*ny*nz product against int overflow, and returns false if any
  expected value is missing.
- read_vdata_palgrid() checks the return value, logs a warning, and
  propagates the failure to its caller.
- Add test_read_cube.cpp covering valid files, truncated data,
  invalid dimensions, negative natom, and failure propagation through
  read_vdata_palgrid().

Verification: cmake --build build --target MODULE_IO_read_cube
MODULE_IO_rho_io; OMP_NUM_THREADS=1 ctest -V -R
"MODULE_IO_read_cube|MODULE_IO_rho_io" (10/10 passed);
agent_governance_check.py --staged (no blocking findings; docs update
not required, no INPUT behavior change).

* Fix: abort run on cube read failure instead of hanging non-root ranks

ModuleIO::read_vdata_palgrid previously returned false on the root rank
when the cube file was missing or malformed, while the other ranks
entered Parallel_Grid::bcast() and blocked in MPI_Recv waiting for data
that would never be sent. Replace the early returns with
ModuleBase::WARNING_QUIT so the run terminates on every rank.

This makes the bool error code meaningless, so remove the now-dead
fallback paths that depended on it: the nspin=2/4 "rearrange electron
density later" branch, the meta-GGA tau TF fallback, the atomic-rho
fallback for failed file reads, and the read_error/read_kin_error
plumbing in Charge::init_rho. Update the unit test to expect death on a
malformed cube file.

* Docs: add 3.10-LTS filename notes for out_hsk and out_hsr in hs_matrix.md

The online documentation for hs_matrix.md only described the new
(develop) filenames for Hamiltonian/overlap matrix output, which
confused LTS users. Add explicit notes mapping out_hsk to the LTS
keyword out_mat_hs (files data-0-H, data-0-S) and out_hsr to the
LTS keyword out_mat_hs2 (files data-HR-sparse_SPIN0.csr,
data-SR-sparse_SPIN0.csr).

Fixes #6458

* Fix: restore init_chg=auto atomic-density fallback for missing charge files

Commit 6601fb6 made a failed cube read abort instead of hanging non-root
ranks, but in doing so it dropped the init_chg=auto semantics: when no
density file exists, auto must silently fall back to the atomic density.
This broke ASE/abacuslite MD, whose first ionic step legitimately runs in a
fresh directory with no charge file and relies on that fallback.

read_rho_file/read_kin_file now probe for the cube file on the parsing rank
and broadcast the result with Parallel_Common::bcast_bool, so all ranks agree
to skip together (no MPI hang) instead of aborting inside read_vdata_palgrid.
init_rho aborts only for init_chg=file; auto falls back to atomic density
(and TF tau for meta-GGA). Add regression tests for both paths.

Verified: cmake --build build --target elecstate passes;
agent_governance_check.py --base HEAD --head HEAD reports no findings.

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
feat: write_bse_ab support MPI parallel.
…8032)

The examples/ tree was renumbered (02_scf, 17_relax, 18_md, 34_gpu and
so on) but 22 links under docs/advanced/ still pointed at the old
directory names and returned 404 on GitHub. Each link now points at
the directory that holds the same case.
* Fix(base): own spherical Bessel root buffer with vector

* Fix(ri): own RPA Coulomb helpers across exceptions
… convention (#8017)

* Split LibXC threshold masks for vrho and vsigma following QE convention

Add cal_sgn_vxc returning separate masks: exc/vrho are evaluated down to
rho_threshold_lda (1e-10), while only the vsigma gradient term is
suppressed below rho_threshold_gga (1e-6) / grho_threshold_gga (1e-10),
matching Quantum ESPRESSO's libxc interface (XClib/xc_wrapper_gga.f90).

* Add unit tests for cal_sgn_vxc threshold masks

Cover the two-tier mask logic in libxc_tools.cpp: vrho kept down to
rho_threshold_vrho while vsigma is suppressed below
rho_threshold_vsigma / grho_threshold_vsigma, GGA vs LDA behavior,
and joint spin-channel masking for nspin=2.

* Test: update 08_EXX HSE reference values for the LibXC threshold masks

The two-tier vrho/vsigma threshold masks in v_xc_libxc shift the
HSE energies, forces and stresses of the H2O-based EXX cases, which
have low-density regions. Regenerate the six affected result.ref
files with the new code (values identical to the CI run of this PR);
totaltimeref entries are kept unchanged.

---------

Co-authored-by: Mohan Chen <mohanchen@pku.edu.cn>
…XX on CPU and GPU (#8018)

* Feature(pw): batched FFTs and small ecut_exx grid for EXX on CPU and GPU

Unify all EXX PW entry points (act_op, act_op_kpar, cal_exx_energy_op)
on one code path built on shared primitives; the batched kernels are a
device specialization of the per-band operations, selected inside the
primitives (batch_active).

- exx_batch kernels templated on Device: host loops + FFTW plan_many on
  CPU (single-precision plans compiled only with ENABLE_FLOAT_FFTW, stubs
  otherwise), CUDA kernels + cuFFT as DEVICE_GPU specializations.
- QE ecutfock-style small FFT grid from ecutexx when every |k+G|^2 fits,
  on CPU and GPU; falls back to the full grid with a warning otherwise.
  The full-grid batched path runs on both devices when the box is local.
- Physics fix: the Fock operator weight now uses the source-state
  occupation f_{mq} and the source k-point weight (was the target-k wg in
  act_op and target-k wk in act_op_kpar). With k-dependent occupations
  (smearing) the old operator was inconsistent with the energy and the
  EXX outer loop never converged.
- The per-(q,m) scalar MPI_Bcast of wg is replaced by one broadcast of
  the occupation row + wk per source k-point via Parallel_Common wrappers.
- stress_exx G-sum truncated to the ecutexx sphere (CPU), consistent with
  the operator and energy.
- Docs: ecutexx describes the small-grid behavior and fallback.

Governance exception: the PARAM/GlobalV budget flags are migration-
neutral moves - the refactor rewrites existing blocks (act_op_kpar,
cal_exx_energy_op, setup) that already read PARAM.inp.nspin/ecutexx and
GlobalV::MY_POOL in the same style as the surrounding module.

Verified against main (4f2a397): build/rel (g++ MPI) and
build_abacus_gnu (CUDA 13.1); 19-case regression matrix
(097_PW_PBE0{,_FM,_COND} x ACE/noACE x full/small grid x CPU/GPU) -
full-grid cases match the pre-change code at 1e-14, CPU small grid
matches GPU small grid to 1e-13, metallic noACE case converges in 3 EXX
outer iterations (did not converge before the wg fix).

* Feature(pw): decouple the EXX band chunking from the grid choice

The small ecut_exx grid and band batching were bound in a single
predicate; they are orthogonal concerns (grid = which FFT box, chunk =
how many bands per batched round). Split them:

- exx_grid_active() answers the grid question (small grid usable, or
  the full box local); exx_band_chunk() answers the band question via
  the new exx_batch_size INPUT (default 0 = all bands, identical to the
  previous behavior; a positive value processes bands in chunks of that
  width with a proportionally smaller work-buffer footprint).
- cache_psi_nk_real, apply_exx_nbatched and the energy pair-density
  loop process bands in chunks; psi_nk_real_cache still holds all bands
  (it is the reuse floor across (iq, m)), only the work buffers shrink.
- Docs: exx_batch_size in parameters.yaml and input-main.md.

Governance exception: exx_band_chunk() reads PARAM.inp.exx_batch_size,
the module's established style for INPUT values (same as ecutexx).

Verified: 24-case regression (097_PW_PBE0{,_FM,_COND} x ACE/noACE x
full/small grid x CPU/GPU, plus exx_batch_size 1/3/5 variants) - the
default keeps every previous value bitwise, chunked runs are bitwise
identical to unchunked (band blocks are disjoint and the (q,m)
accumulation order is unchanged).

* Refactor(pw): pass EXX PW configuration through General_Exx_Info instead of globals

The EXX PW operator and stress read their configuration (nspin, ecutexx,
exx_batch_size, exxace, exx_gamma_extrapolation) and the MPI layout
(KPAR, MY_RANK, MY_POOL) directly from PARAM/GlobalV. Inject them
explicitly instead:

- General_Exx_Info carries the PW EXX INPUT values, resolved once in
  init_general_exx_info: exxace, gamma_extrapolation, exx_batch_size,
  and ecut_exx (ecutexx when set, else ecutrho) with the user-set flag.
- OperatorEXXPW takes the General_Exx_Info plus the runtime values
  (nspin, kpar, my_rank, my_pool) at construction and snapshots them as
  members; all internal PARAM/GlobalV reads are gone.
- Stress_PW::stress_exx takes the General_Exx_Info in place of the
  separate hybrid_alpha/coulomb_param parameters.

Governance: the global-dependency budget of the PR goes from +12 to
-28 (added=4, removed=32); the remaining 4 added references are the
single snapshot point in HamiltPW, the orchestration layer that owns
these values.

Verified: build/rel (g++ MPI) rebuilds; tests/01_PW/097_PW_PBE0 matches
result.ref (etot to 1.8e-14 eV, stress sum exact at 1e-6 kbar);
exx_batch_size=1 is bitwise identical to unchunked; ecutexx=20 engages
the small FFT grid at 1 rank and prints the distributed-box fallback
warning once at 2 ranks, both converging to the same energy (2.4e-14
eV).

* Docs: quote exx_batch_size default_value in parameters.yaml

Match the --generate-parameters-yaml output exactly; the documentation
consistency CI check requires docs/parameters.yaml to be byte-identical
to the generated file.

* Fix(build): wire EXX batch kernels into the legacy Makefile build

The batched-EXX commits added kernels/exx_batch_op.cpp,
kernels/exx_batch_op_float.cpp (ENABLE_FLOAT_FFTW) and its stub, plus
kernels/cuda/exx_batch_op.cu, and wired them into CMake only. The
Makefile build then failed at link time with undefined references to
hamilt::exx_batch_* from the rewritten op_pw_exx.cpp.

Add exx_batch_op.o and exx_batch_op_float_stub.o to OBJS_HAMILT; the
legacy Makefile has no float-FFTW switch, matching the CMake
ENABLE_FLOAT_FFTW=OFF branch. VPATH already covers the kernels dir.

* Fix(pw): reject non-smaller EXX FFT grids, clamp ecutexx to ecutrho

Addresses the Copilot review on #8018: setup_exx_small_grid only
rejected a box exactly equal to the wavefunction box, so ecutexx >
ecutrho (or FFT-friendly dimension rounding) enabled the small-grid
path with sg_nxyz > wfcpw->nrxx, overflowing psi_mq_real/psi_nk_real.
Compare box volumes against wfcpw->nrxx instead of dimensions.

Also clamp a user-set ecutexx to ecutrho in init_general_exx_info:
the pair density carries no G-components beyond the ecutrho sphere,
while the EXX buffers (density_recip, pot, ...) and the rhopw_dev box
are sized by ecutrho, so larger values only corrupt memory. This
overflow predates this PR (same pattern on develop); the clamp turns
it into a warning plus full-grid fallback.

Docs: no update needed; the existing ecutexx description already
documents the full-grid fallback for a non-smaller box.

Verification (GNU+CUDA build, OMP_NUM_THREADS=1, tests/01_PW/097_PW_PBE0):
- default: E_TOT -30.5936375896797834 eV, bit-identical to result.ref
- ecutexx=100 (> ecutrho=40): ran clean with the clamp warning and
  full-grid fallback, E_TOT bit-identical to default; valgrind shows
  no invalid accesses in ABACUS code (previously heap corruption and
  SIGSEGV/SIGABRT)
- ecutexx=15: small grid (9,9,9) vs (15,15,15) engaged as before,
  dE = 2.7e-5 eV from the intended pair-density truncation
…SdftPW to StoHamiltPW (#8012)

* Move HSolverPW_SDFT into module_stodft as HSolverSdftPW

After #7974 inverted the hamilt/hsolver dependency through the
HSOperator/HSMatrix interfaces, HSolverPW_SDFT was the only thing left in
source_hsolver still reaching into source_hamilt, and the only thing left
reaching into source_pw. It also closed the one real include cycle in the
area:

    sto_elecond.h -> source_hsolver/hsolver_pw_sdft.h
                  -> source_pw/module_stodft/{hamilt_sdft_pw.h, sto_iter.h}

It is not a solver algorithm: it owns a Stochastic_Iter, takes a
HamiltSdftPW*, and orchestrates the stochastic-DFT SCF step. It belongs in
module_stodft, so move it there. Converting it to HSOperator instead would
have dropped the hamilt include but left the source_pw edges and the cycle.

Renamed to match the sibling it now sits next to, which is the same class
of thing for the Hamiltonian side:

    hamilt_sdft_pw.h  -> hamilt::HamiltSdftPW  : public HamiltPW
    hsolver_sdft_pw.h -> hsolver::HSolverSdftPW : public HSolverPW

That also drops the CamelCase/SCREAMING_SNAKE mix in HSolverPW_SDFT and
puts the basis suffix last, as every other name in that directory does. The
namespace stays hsolver: the class implements the HSolver role, and the
directory already holds hamilt::HamiltSdftPW, so "namespace = role,
directory = feature area" is the established rule here.

Result, measured on the non-test sources:

    source_hsolver -> source_hamilt   1 file -> 0
    source_hsolver -> source_pw       2 files -> 0

The cycle is gone. What remains between the two is source_pw ->
source_hsolver (hsolver_pw.h, para_lin_tf.h), which is the correct
direction.

The unit test moves with the code and keeps __MPI via
KEEP_FEATURE_DEFINITIONS, because module_stodft/test disables it and this
test calls MPI_Init and has mocks taking MPI_Comm unconditionally. Its
CTest target is renamed MODULE_HSOLVER_sdft -> MODULE_PW_Sto_HSolver_UTs so
it groups with its new module and its two neighbours.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Rename HamiltSdftPW to StoHamiltPW and move it out of namespace hamilt

module_stodft names its files sto_* and keeps its classes (StoChe,
Sto_DOS, Stochastic_Iter, ...) in the global namespace.
hamilt_sdft_pw.{h,cpp} is the leftover from when this code lived under
hamilt_stodft, so bring it in line:

  hamilt_sdft_pw.{h,cpp}  -> sto_hamilt_pw.{h,cpp}
  hamilt::HamiltSdftPW    -> StoHamiltPW (global namespace)
  test_hamilt_sto.cpp     -> test_sto_hamilt_pw.cpp

Names that were found through the enclosing namespace (HamiltPW,
Operator, hpsi_norm_op) are now qualified with hamilt::. The timer and
classname labels follow the class name. No logic change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Rename HSolverSdftPW to StoHSolverPW and move it out of namespace hsolver

Same treatment as StoHamiltPW in the previous commit, so both classes
this PR brings into module_stodft follow the directory's sto_* file
naming and global-namespace convention:

  hsolver_sdft_pw.{h,cpp}   -> sto_hsolver_pw.{h,cpp}
  hsolver::HSolverSdftPW    -> StoHSolverPW (global namespace)
  test_hsolver_sdft_pw.cpp  -> test_sto_hsolver_pw.cpp

The base class is now spelled hsolver::HSolverPW. The TITLE/timer
labels and the local object in ESolver_SDFT_PW follow the class name.
The test's hsolver:: mocks stay in namespace hsolver since they mock
the base-class side. No logic change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Qualify HamiltPW in the StoHamiltPW mock constructor of test_sto_tool

The mock constructor's initializer list named the base as HamiltPW,
which was only found while the class lived in namespace hamilt. Spell
it hamilt::HamiltPW, as sto_hamilt_pw.cpp already does. Fixes the
MODULE_PW_Sto_Tool_UTs build failure in the Test and CUDA Test jobs.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Disable __CUDA/__ROCM for the module_stodft unit tests

source_hsolver/test, where the SDFT solver test used to live, builds
without __CUDA/__ROCM. module_stodft/test did not, so after the move the
CUDA build compiled the DEVICE_GPU instantiations of StoHSolverPW,
HSolverPW and FFT_CUDA into MODULE_PW_Sto_HSolver_UTs without linking
their GPU implementations, and the link failed.

Disable both for the directory, as source_hsolver/test does. All three
tests here only exercise CPU code; for Sto_Tool and Sto_Hamilt this just
drops GPU template instantiations they never call.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* change out_interval to out_freq_ion

* Clean up remaining out_interval references

- Remove unused out_interval variable declaration from input_parameter.h
- Replace out_interval with out_freq_ion in example INPUT files
- Add missing cmath include for std::pow in input_parameter.h

* feat: always output magnetic moments in print_stru_file for nspin=2/4

- Element line #magnetism now outputs real initial value instead of 0.0000
  - nspin=2: atoms[it].mag[0]
  - nspin=4: |atoms[it].m_loc_[0]|
- Add comment "(default, overridden by per-atom mag below)" to clarify
  element-line value is a default that can be overridden
- Per-atom mag is now always printed for nspin=2/4:
  - magmom=true && atom_mulliken non-empty: Mulliken analysis values
  - otherwise: initial values from atoms[it].mag[ia] / m_loc_[ia]
- Update unit tests in unitcell_test.cpp:
  - Fix existing expectations for nspin=2 without magmom
  - Add new tests for non-zero initial mag, nspin=4 initial mag,
    and nspin=4 Mulliken mag

Also update test case relax_out_hk_spin2:
- STRU: add mag 1.0 / mag -1.0 for antiferromagnetic initial state
- INPUT: symmetry 1 -> 0 to allow AFM structure

* feat: extend out_stru to scf/nscf, opt-in only

out_stru was previously only effective for relax/cell-relax; scf and
nscf runs never wrote structure files (and nscf even force-reset
out_stru to 0).

Now:
- Relax_Driver::stru_out/final_out accept scf/nscf in addition to
  relax/cell-relax. For scf/nscf only STRU_FINAL (or STRU_FINAL.cif) is
  written, since the structure is identical to the single step; the
  per-step STRU_NOW and numbered STRU{istep} files remain relax-only,
  as do the relax-specific screen messages.
- out_stru reset_value only forces 0 when the user did not set it
  explicitly (item.is_read()); scf is added to the offlist so that scf
  and nscf produce no structure output by default and become opt-in.
  relax/cell-relax keep their default value 1, unchanged.
- Docs (parameters.yaml, input-main.md, INPUT description) updated to
  describe the new effective scope and scf/nscf opt-in behavior.
- read_input_item_test.cpp adapted: OutStru test now covers both
  "reset when not read" and "preserve explicit user value" cases.

Verification: code edited only; build/tests not run (per user request,
awaiting approval).

* test: switch 02_NAO_Gamma/relax_bfgs2 to L-BFGS as relax_lbfgs

The bfgs 2 relax path is already covered by relax_out_hk_spin2 in the
same directory and by tests/03_NAO_multik/relax_bfgs2 with multiple
k-points. Repurpose the Gamma-only case to exercise relax_method=lbfgs,
which had no integration-test coverage, and rename the directory
accordingly. result.ref is intentionally dropped; regenerate it with
run_debug.sh ref before relying on this test.

* feat: add force output to print_stru_file in Angstrom/eV units

- Add const ModuleBase::matrix& force parameter to print_stru_file
- When force is provided (nr==nat, nc==3), output positions in Angstrom
  and forces in eV/Angstrom with 'f' keyword
- Update unit note in ATOMIC_POSITIONS header accordingly
- Update all existing test calls to pass empty matrix for backward
  compatibility

* feat: pass force matrix to stru_out and final_out for STRU force output

- Add force parameter to stru_out() and final_out()
- Move force declaration outside while loop so final_out can access it
- Pass force to print_stru_file for STRU_NOW, STRU{istep+1}, STRU_FINAL

* test: add dedicated print_cell tests and fix Angstrom force output units

- Add test_print_cell.cpp with MODULE_CELL_print_cell_test target covering
  nspin=1/2/4 magnetic moment output and force output
- Move PrintSTRU tests out of unitcell_test.cpp into the new test file
- When forces are present, always emit Cartesian_angstrom coordinates
  (tau * lat0 * BOHR_TO_A) and forces in eV/Angstrom; internal tau was
  previously treated as Bohr and missed the lat0 factor, and direct
  fractional coordinates were incorrectly converted as Cartesian
- Include source_base/matrix.h in print_cell.h for the default force
  argument's complete type

* feat: emit STRU in Angstrom throughout and use single-space coordinates

- Always print LATTICE_CONSTANT as 1 Angstrom in Bohr and write lattice
  vectors directly in Angstrom (lat0 * latvec * BOHR_TO_A), keeping the
  STRU self-consistent when read back as Cartesian_angstrom
- Always use Cartesian_angstrom for Cartesian positions; Direct is only
  emitted for fractional coordinates without forces, and forces force
  Cartesian_angstrom output
- Print position, velocity and force components separated by single
  spaces instead of fixed-width columns
- Update print_cell unit tests for the new headers and spacing

* style: print STRU forces with 6 decimal places

Reduce force output precision from 10 to 6 decimals in eV/Angstrom;
positions and velocities remain at 10 decimals.

* update relax_out_hk_spin2/STRU

* update reference data in 02_NAO_Gamma

* style: single space before inline comments in STRU output

Normalize STRU_NOW formatting so values are followed by exactly one
space before the trailing comment, and per-atom mag uses one space
after the mag keyword. Update print_cell test expectations (including
stale lat0/coordinate values) to match current BOHR_TO_A output.

---------

Co-authored-by: abacus_fixer <mohanchen@pku.eud.cn>
   Remove the five AbacusDevice_t-keyed memory wrappers and their tests;
   they have no call site in the tree or in the history.

   Refs #8039
@Cstandardlib
Cstandardlib deleted the fix/remove-dead-device-memory-wrappers branch September 28, 2026 09:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.