Search arXivSearch

arXiv subjects

Ziang Yan

Publications and source records attributed to Ziang Yan.

At least 19 recordsLinked to original sources

Anisotropic redshift distributions in photometric galaxy clustering and their cosmological impact

Photometric galaxy clustering is a major probe of large-scale structure whose interpretation relies on accurate characterization of the galaxy redshift distribution, $n(z)$. Standard analyses assume the $n(z)$ of a selected lens sample to be isotropic, whereas observational systematics can imprint spatial variations. We quantify how this anisotropy biases angular galaxy clustering, internal consistency, and cosmological inference. We develop an analytical formalism for clustering under spatially varying $n(z)$ and forward-model realistic observational systematics, galaxy selection, and tomographic binning. For a fiducial magnitude-limited sample, modeling $w(θ)$ with a single global $\bar n(z)$ underestimates the clustering signal by up to $\sim10\%$ at small scales in the highest-redshift bin, decreasing towards lower redshift. In a $2\times2$pt analysis, cosmological parameter shifts remain below $1σ$ for KiDS-like and DES-like setups. For an LSST-like setup retaining KiDS-like selection and observing-condition variation, $σ_8$ is biased low by $1.7σ$ and $Ω_m$ high by $1.4σ$, while the high-redshift galaxy bias shifts by up to $\sim4σ$. The anisotropy also affects whether the probes are consistent enough to combine: a calibrated posterior-predictive internal-consistency test finds the contaminated angular clustering improbable given the joint cosmic-shear and galaxy--galaxy-lensing constraints (median calibrated $\tilde p\approx8\times10^{-4}$ under LSST-like noise). Anisotropic redshift distributions are therefore a potentially significant systematic for future photometric clustering analyses. Our forward model can be used to mitigate it given a realistic selection function and galaxy population model.

astro-ph.CO

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse video types, making them effective only in specific domains. High computational demands further restrict their efficiency and scalability. Moreover, most models are only partially open, with key components such as training code, strategy, or datasets unavailable, which hinders reproducibility and slows community-driven development. To address these issues, we introduce VideoChat3, a fully open, efficient, and generalist video-centric MLLM. VideoChat3 advances video understanding through two complementary designs. For efficiency, we introduce Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for Streaming Video Perception, which enables efficient spatiotemporal representation and reduces the cost of processing video inputs during training and inference. For effectiveness, we develop a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets: VideoChat3-Academic2M, VideoChat3-LV116K, and VideoChat3-OL617K, covering general, long-form, and streaming video scenarios, improving the model's generalization across domains. By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.

cs.CV

Intern-S2-Preview: Scientific Agentic Foundation Model

Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across long task horizons. We present Intern-S2-Preview, a series of scientific agentic foundation models designed to support multimodal scientific understanding, reasoning, generation, and long-horizon tasks. The training pipeline begins with scientific multimodal pre-training over rendered scientific documents, interleaved image-text data, and diverse scientific corpora. Starting from the pretrained checkpoint, we apply a unified post-training pipeline consisting of supervised fine-tuning, scalable multi-task reinforcement learning (RL), black- and white-box agentic RL, and on-policy distillation. This pipeline is supported by practical techniques that improve rollout and training stability and efficiency, including partial rollout with off-policy correction, adaptive length regularization, online speculative decoding, robust multi-task optimization, and trace-aware experience assembly for agentic tasks. At the architecture level, Intern-S2-Preview-397B extends time series modelling from efficient long-sequence understanding to numerical forecasting, while Memory Decoder is studied as a separate memory-augmented path for rapid scientific specialization without modifying the frozen 397B backbone. Evaluations across scientific, multimodal, agentic, and general-purpose benchmarks show that Intern-S2-Preview-397B achieves competitive or leading results in multiple settings. The time series modules improve scientific signal understanding and forecasting on SciTS, while the separate Intern-MemDec-4B extension improves the Biology-Instructions average score from 56.92 to 60.32 without modifying the frozen 397B backbone.

cs.LG

EchoVLA: Robotic Vision-Language-Action Model with Synergistic Declarative Memory for Mobile Manipulation

Recent progress in Vision-Language-Action (VLA) models has enabled embodied agents to interpret multimodal instructions and perform complex tasks. However, existing VLAs are mostly confined to short-horizon, table-top manipulation, lacking the memory and reasoning capability required for mobile manipulation, where agents must coordinate navigation and manipulation under changing spatial contexts. In this work, we present EchoVLA, a memory-aware VLA model for mobile manipulation. EchoVLA incorporates a synergistic declarative memory inspired by the human brain, consisting of a scene memory that maintains a collection of spatial-semantic maps and an episodic memory that stores task-level experiences with multimodal contextual features. The two memories are individually stored, updated, and retrieved based on current observations, task history, and instructions, and their retrieved representations are fused via coarse- and fine-grained attention to guide base-arm diffusion policies. To support large-scale training, we further introduce MoMani, an automated benchmark that generates expert-level trajectories through multimodal large language model (MLLM)-guided planning and feedback-driven refinement, supplemented with real-robot demonstrations. Comprehensive simulated and real-world results demonstrate that EchoVLA substantially improves overall performance, e.g., it achieves the highest success rates of 0.52 on manipulation/navigation tasks and 0.31 on mobile manipulation tasks in simulation, exceeding the strong baseline $π_{0.5}$ by +0.20 and +0.11, respectively.

cs.RO

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs

Video multimodal large language models (MLLMs) can describe what happens in a video, but rarely identify when the supporting evidence occurs. We study generalist video temporal grounding, in which one model predicts a variable-cardinality set of evidence intervals across video lengths, domains, query forms, and viewpoints. Existing training strategies are misaligned with this set-valued task: long-video labels often rely on brittle one-pass annotation, while reinforcement-learning rewards either fail to distinguish non-overlapping predictions or require fragile segment matching. TimeLens2 treats temporal evidence as an interval set throughout supervision and optimization. TimeLens2-93K constructs reliable multi-span supervision through caption-derived proposals, independent localization, cross-agent consensus, semantic verification, and boundary refinement. Our temporal Wasserstein reward computes exact one-dimensional \(W_1\) between uniform distributions over merged interval supports, providing dense, matching-free feedback under unequal cardinalities and equivalent fragmentation; temporal IoU complements it with precise-overlap feedback. Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters. The 2B, 4B, and 8B variants improve over their Qwen3-VL backbones by 14.2, 13.0, and 18.1 mIoU points, respectively.

cs.CV

KiDS-Legacy: The consistency test of the large-scale structure with Bernardeau-Nishimichi-Taruya transform

We perform the first $k$-cut cosmic shear analysis of the KiDS-Legacy survey. This method uses the Bernardeau-Nishimichi-Taruya (BNT) transform to construct weak-lensing kernels that are more localised than conventional ones, and remove information from selected physical scales while retaining the constraining power of the targeted range. Removing the scale of $k \geq 0.33~\mathrm{Mpc}^{-1}$ from the KiDS-Legacy pseudo-$C_\ell$ data vector, and using a covariance matrix whose Gaussian component is computed from the theoretical data vector, we find $S_8 = 0.798 \pm 0.045$. This agrees with both the fiducial KiDS-Legacy bandpower result and our no-$k$-cut pseudo-$C_\ell$ posterior to within $0.1σ$, indicating no significant bias from nonlinear astrophysical feedback at the precision of KiDS-Legacy. We also study the case in which the Gaussian covariance is computed from the observed data vector. In this setup, the same scale cut of $k < 0.33~\mathrm{Mpc}^{-1}$ gives a much lower $S_8=0.717_{-0.046}^{+0.047}$. Further $k$-cut tests reveal a mild scale-dependent trend, with larger physical scales preferring lower $S_8$ values and a maximum low- versus high-$k$ deviation of $1.80σ$. Mock tests show that this behaviour is not produced by the covariance prescription or data vector alone, but may arise from their interplay. These results show that BNT $k$-cuts provide both a mitigation strategy for nonlinear systematics and a diagnostic of weak-lensing inference pipelines.

astro-ph.CO

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning

Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts largely focus on text-dominant settings, leaving long-horizon multimodal tasks underexplored. This gap is evident in video tasks requiring sustained temporal understanding and iterative interaction. We present InternVideo3, a framework enhancing these capabilities via Multimodal Contextual Reasoning (MCR). MCR treats understanding as a closed-loop process over a shared, evolving context containing observations, instructions, reasoning, tool actions, and memory. This frames long-video understanding as evidence accumulation and verification. To ensure efficiency, we introduce Multimodal Multi-head Latent Attention (M^2LA), a token-preserving reparameterization compressing KV-cache states while retaining the full token stream. Our staged training includes continued pretraining, short-to-long supervised fine-tuning, rule-based reinforcement learning, and on-policy distillation. Experiments show InternVideo3 achieves strong performance on benchmarks like Video-MME, MLVU, and EgoSchema. We further instantiate the model as a video agent with retrieval tools, demonstrating robust evidence-grounded behavior. Our results suggest that efficient context handling and closed-loop reasoning are vital for adapting open multimodal models toward long-horizon visually grounded agency.

cs.CV

ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation

On-policy distillation (OPD) improves reasoning by training a student on trajectories sampled from its own policy under supervision from a teacher. In multimodal reasoning, a common extension is to use a privileged teacher that observes training-time-only signals such as reference answers or rationales. However, such answer-side privilege creates a train-test mismatch: the teacher's supervision may depend on signals unavailable to the student, encouraging shortcut imitation rather than visually grounded reasoning. We propose ViCuR, a visually grounded privileged-teacher distillation framework that replaces answer-side privilege with visual cues (query-related evidence in the input). Because these cues are derived from the same visual input available at inference, their evidence is recoverable by the student. To support this, ViCuR introduces a lightweight cue recovery module that uses dedicated sink-token cross-attention during prefill to aggregate task-relevant visual evidence into an internal representation, without changing the inference interface or requiring auxiliary cue-generation losses. Across seven benchmarks with Qwen3-VL-2B and 8B students, ViCuR consistently improves over answer-based on-policy self-distillation by +1.19 and +1.24 on overall average performance. It also extends naturally to stronger-teacher OPD, surpassing OPD baselines by +0.64 and +1.08, with consistent out-of-domain gains at the 8B scale. These results show that, in multimodal on-policy distillation, the design of teacher privilege is as important as teacher strength.

cs.CV

Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction

Video event prediction (VEP) requires models to infer unobserved future states from partial video evidence. Existing video MLLMs usually verbalize intermediate future reasoning in text space: once visual evidence is verbalized, fine-grained motion, geometry, and interaction cues can be lost, leading to plausible but visually ungrounded hallucinations. We introduce Future-L1, an interleaved latent visual reasoning framework that lets an MLLM alternate between language tokens and continuous latent visual spans during autoregressive decoding. To train this capability, we construct Future-L1-50K by selecting examples where future visual hints help prediction and align latent states to future-frame embeddings, then further optimize sampled latent trajectories with LA-DAPO, a latent-aware RL objective with outcome-contrastive and temporal-diversity rewards. Future-L1 achieves new state-of-the-art results on both benchmarks: on FutureBench, it improves Qwen3-VL-8B from 61.0 to 85.4 and exceeds the previous best Video-CoE by 10.4 points; on TwiFF-Bench, it improves the average score from 2.44 to 3.04. These results suggest that future-oriented video reasoning benefits from preserving intermediate visual semantics in latent space rather than translating every reasoning step into text.

cs.CV

FreeRet: MLLMs as Training-Free Retrievers

Multimodal large language models (MLLMs) are emerging as versatile foundations for mixed-modality retrieval. Yet, they often require heavy post-hoc training to convert them into contrastive encoders for retrieval. This work asks: Can off-the-shelf MLLMs serve as powerful retrievers without additional training? We present FreeRet, a plug-and-play framework that turns any MLLM into a two-stage retriever. FreeRet first derives semantically grounded embeddings directly from the model for fast candidate search, and then exploits its reasoning ability for precise reranking. The framework contributes three advances: bypassing lexical alignment layers to obtain semantically faithful embeddings, conditioning representation generation with explicit priors, and mitigating framing effect in reranking via neutral choice framing. On the MMEB and MMEB-V2 benchmarks spanning 46 datasets, FreeRet substantially outperforms models trained on millions of pairs. Beyond benchmarks, FreeRet is model-agnostic and scales seamlessly across MLLM families and sizes, preserves their generative abilities, supports arbitrary modality combinations, and unifies retrieval, reranking, and generation into end-to-end RAG within a single model. Our findings demonstrate that pretrained MLLMs, when carefully harnessed, can serve as strong retrieval engines without training, closing a critical gap in their role as generalists.

cs.CV

Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Existing multimodal large language models for long-video understanding predominantly rely on uniform sampling and single-turn inference, limiting their ability to identify sparse yet critical evidence amid extensive redundancy. We introduce Video-o3, a novel framework that supports iterative discovery of salient visual clues, fine-grained inspection of key segments, and adaptive termination once sufficient evidence is acquired. Technically, we address two core challenges in interleaved tool invocation. First, to mitigate attention dispersion induced by the heterogeneity of reasoning and tool-calling, we propose Task-Decoupled Attention Masking, which isolates per-step concentration while preserving shared global context. Second, to control context length growth in multi-turn interactions, we introduce a Verifiable Trajectory-Guided Reward that balances exploration coverage with reasoning efficiency. To support training at scale, we further develop a data synthesis pipeline and construct Seeker-173K, comprising 173K high-quality tool-interaction trajectories for effective supervised and reinforcement learning. Extensive experiments show that Video-o3 substantially outperforms state-of-the-art methods, achieving 72.1% accuracy on MLVU and 46.5% on Video-Holmes. These results demonstrate Video-o3's strong multi-hop evidence-seeking and reasoning capabilities, and validate the effectiveness of native tool invocation in long-video scenarios.

cs.CV

KiDS-Legacy: Constraining dark energy, neutrino mass, and curvature

We constrained minimally extended cosmological models with the cosmic shear analysis of the final data release from the Kilo-Degree Survey (KiDS-Legacy) in combination with external probes. Due to the consistency of the KiDS-Legacy analysis with the cosmic microwave background (CMB), we could combine these datasets reliably for the first time. Additionally, we used CMB lensing, galaxy redshift-space distortions, and baryon acoustic oscillations. We assessed, in turn, the effects of spatial curvature, varying neutrino masses, and an evolving dark energy component on cosmological constraints from KiDS-Legacy alone and from KiDS-Legacy combined with external probes. We find KiDS-Legacy to be consistent with the fiducial flat $Λ$-cold dark matter ($Λ$CDM) analysis with $c^2 \sum m_ν\leq 1.5\,$eV, $w_0 = -1.0\pm 0.7$, and $w_a = -1.3^{+1.9}_{-2.0}$ while $Ω_K = 0.08^{+0.16}_{-0.17}$ (1$σ$ bounds) with an almost equal goodness of fit. The $w_0w_a$CDM model is not a significant improvement over $Λ$CDM when cosmic shear and CMB lensing are combined, yielding a Bayes factor $B = 0.07$. If all probes are combined, however, $B$ increases to 2.73, corresponding to a $2.6σ$ suspiciousness tension. The constraint on $S_8 = σ_8\sqrt{Ω_\mathrm{m}/0.3}$ is robust to opening up the parameter space for cosmic shear. Adding all external datasets to KiDS-Legacy, we find $S_8 = 0.816 \pm 0.006$ in $Λ$CDM and $S_8 = 0.837 \pm 0.008$ in $w_0 w_a$CDM for all probes combined.

astro-ph.CO

Improved photometric redshift estimations through self-organising map-based data augmentation

We introduce a framework for the enhanced estimation of photometric redshifts using Self-Organising Maps (SOMs). Our method projects galaxy Spectral Energy Distributions (SEDs) onto a two-dimensional map, identifying regions that are sparsely sampled by existing spectroscopic observations. These under-sampled areas are then augmented with simulated galaxies, yielding a more representative spectroscopic training dataset. To assess the efficacy of this SOM-based data augmentation in the context of the forthcoming Legacy Survey of Space and Time (LSST), we employ mock galaxy catalogues from the OpenUniverse2024 project and generate synthetic datasets that mimic the expected photometric selections of LSST after one (Y1) and ten (Y10) years of observation. We construct 501 degraded realisations by sampling galaxy colours, magnitudes, redshifts and spectroscopic success rates, in order to emulate the compilation of a wide array of realistic spectroscopic surveys. Augmenting the degraded mock datasets with simulated galaxies from the independent CosmoDC2 catalogues has markedly improved the performance of our photometric redshift estimates compared to models lacking this augmentation, particularly for high-redshift galaxies ($z_\mathrm{true} \gtrsim 1.5$). This improvement is manifested in notably reduced systematic biases and a decrease in catastrophic failures by up to approximately a factor of 2, along with a reduction in information loss in the conditional density estimations. These results underscore the effectiveness of SOM-based augmentation in refining photometric redshift estimation, thereby enabling more robust analyses in cosmology and astrophysics for the NSF-DOE Vera C. Rubin Observatory.

astro-ph.GA

Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale

We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. Simultaneously, its scientific expertise has been vastly expanded to master over 100 specialized tasks across critical science fields, including chemistry, materials, life sciences, and earth sciences. Achieving this massive scale is made possible by the robust infrastructure support of XTuner and LMDeploy, which facilitates highly efficient Reinforcement Learning (RL) training at the 1-trillion parameter level while ensuring strict precision consistency between training and inference. By seamlessly integrating these advancements, Intern-S1-Pro further fortifies the fusion of general and specialized intelligence, working as a Specializable Generalist, demonstrating its position in the top tier of open-source models for general capabilities, while outperforming proprietary models in the depth of specialized scientific tasks.

cs.LG

InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision

Large-scale video-text pretraining achieves strong performance but depends on noisy, synthetic captions with limited semantic coverage, often overlooking implicit world knowledge such as object motion, 3D geometry, and physical cues. In contrast, masked video modeling (MVM) directly exploits spatiotemporal structures but trails text-supervised methods on general tasks. We find this gap arises from overlooked architectural issues: pixel-level reconstruction struggles with convergence and its low-level requirement often conflicts with semantics, while latent prediction often encourages shortcut learning. To address these, we disentangle the traditional encoder-decoder design into an Encoder-Predictor-Decoder (EPD) framework, where the predictor acts as a latent world model, and propose InternVideo-Next, a two-stage pretraining scheme that builds a semantically consistent yet detail-preserving latent space for this world model. First, conventional linear decoder in pixel MVM enforces the predictor output latent to be linearly projected to, thus separable in pixel space, causing the conflict with semantic abstraction. Our Stage 1 proposes a conditional diffusion decoder and injects reliable image-level semantic priors to enhance semantics and convergence, thus bridging pixel-level fidelity with high-level semantic abstraction. Stage 2 further learns world knowledge by predicting frozen Stage 1 targets within this space, mitigating shortcut learning. Trained on public, unlabeled videos, InternVideo-Next achieves state-of-the-art results across benchmarks and provides a scalable path toward general video representation learning.

cs.CV

KiDS-Legacy: Constraints on Horndeski gravity from weak lensing combined with galaxy clustering and cosmic microwave background anisotropies

We present constraints on modified gravity from a cosmic shear analysis of the final data release of the Kilo-Degree Survey (KiDS-Legacy) in combination with DESI measurements of baryon acoustic oscillations, eBOSS observations of redshift space distortions, and cosmic microwave background anisotropies from Planck. We study the Horndeski class of modified gravity models in an effective field theory framework employing a parameterisation that satisfies stability conditions by construction and, for the first time, present a cosmological analysis in this inherently stable parameter basis. Cosmic shear constrains the Horndeski parameter space significantly, matching or surpassing the CMB contribution. Adopting the de-mixed kinetic term of the scalar field perturbation, $D_{\rm kin}$, and the deviation of the Planck mass from its fiducial value, $ΔM_*^2\equiv M_*^2-1$, as model parameters, we constrain their present values to be $Δ\hat{M}_*^2=0.32^{+0.07}_{-0.21}$ and $\hat{D}_{\rm kin} = 3.74^{+0.69}_{-1.92}$, which deviate from general relativity at $1.5σ$ and $1.9σ$, respectively. We derive constraints on the structure growth parameter $S_8=0.813^{+0.008}_{-0.011}$, which is compatible with the $Λ$CDM constraint at $0.54σ$. We obtain the deviation of the effective Newtonian coupling from the GR value as $Δμ_{\infty,{\rm eff}}=0.066\pm0.023$, corresponding to a $2.9σ$ significance. Although modified gravity provides a slightly better fit to the data, a model comparison shows only a weak preference for modified gravity at the $1.4σ$ level. When adopting a dynamical dark energy model of the background cosmology, the inferred modified gravity parameter constraints are stable with respect to a $Λ$CDM background, while a mild preference at $1.57σ$ for dynamical dark energy remains.

astro-ph.CO

Redshift Assessment Infrastructure Layers (RAIL): Rubin-era photometric redshift stress-testing and at-scale production

Virtually all extragalactic use cases of the Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) require the use of galaxy redshift information, yet the vast majority of its sample of tens of billions of galaxies will lack high-fidelity spectroscopic measurements thereof, instead relying on photometric redshifts (photo-$z$) subject to systematic imprecision and inaccuracy best encapsulated by photo-$z$ probability density functions (PDFs). We present the version 1 release of Redshift Assessment Infrastructure Layers (RAIL), an open source Python library for at-scale probabilistic photo-$z$ estimation, initiated by the LSST Dark Energy Science Collaboration (DESC) with contributions from the LSST Interdisciplinary Network for Collaboration and Computing (LINCC) Frameworks team. RAIL's three subpackages provide modular tools for end-to-end stress-testing, including a forward modeling suite to generate realistically complex photometry, a unified API for estimating per-galaxy and ensemble redshift PDFs by an extensible set of algorithms, and built-in metrics of both photo-$z$ PDFs and point estimates. RAIL serves as a flexible toolkit enabling the derivation and optimization of photo-$z$ data products at scale for a variety of science goals and is not specific to LSST data. We thus describe to the extragalactic science community, including and beyond Rubin the design and functionality of the RAIL software library so that any researcher may have access to its wide array of photo-$z$ characterization and assessment tools.

astro-ph.IM

KiDS-Legacy: WIMP dark matter constraints from the cross-correlation of weak lensing and Fermi-LAT gamma rays

Dark matter dominates the matter content of the Universe, and its properties can be constrained through large-scale structure probes such as the cross-correlation between the unresolved gamma-ray background (UGRB) and weak gravitational lensing. We analysed 15 years of Fermi-LAT data, constructing UGRB intensity maps in ten energy bins (0.5-1000 GeV), and cross-correlated them with KiDS-Legacy shear in six tomographic bins. The measurements were performed using angular power spectra estimated with the pseudo-$C_\ell$ method. No significant cross-correlation is found. Based on this non-detection, we present 95% upper bounds on the weakly interacting massive particle (WIMP) decay rate $Γ_{\rm dec}$ and velocity-averaged annihilation cross-section $\langleσ_{\rm ann} v\rangle$ as functions of mass. We compare our results with bounds from other cosmological tracers and from local probes, and found them to be complementary, particularly at low masses ($\rm GeV/TeV$). In addition, using a Euclid-like lensing survey cross-correlated with Fermi-LAT, we forecast $\sim$2 times tighter limits, highlighting the potential of forthcoming data to strengthen constraints on dark matter annihilation and decay.

astro-ph.CO