Deprecate and remove graph harvesting (#8579)
Deprecate and remove graph harvesting as part of #8489
• Remove the graph_harvesting config option; supplying it (any value,
including false / null ) now raises DeepSpeedConfigError , consistent
with other removed features.
• Remove graph_process / graph_cache and the use_graph argument from
get_global_norm_of_tensors / clip_tensors_by_global_norm , collapsing
them to the direct (non-graph) path.
• Optimize global-norm computation: previously each tensor's p‑th power
was accumulated in-place into a reused, cached scratch buffer
( graph_cache['norm_tensors_compute_buffer'] ) so memory addresses
stayed fixed for graph replay. It now computes each tensor's norm into a
fresh list, stacks them, raises to the p‑th power, and sums after the
model/expert-parallel all-reduce).
• Drop the graph_harvesting plumbing in BF16_Optimizer ,
DeepSpeedEngine , and constants.py ; simplify update_hp_grads (the
graph-only CPU flag-fixup loop was redundant).
• Add a config-rejection test.
---------
Signed-off-by: Hongwei Chen <hongweichen@microsoft.com>