Search arXivSearch

arXiv subjects

Zhihao Zhou

Publications and source records attributed to Zhihao Zhou.

At least 19 recordsLinked to original sources

Intrinsic Limitations of Single Layer Polychromatic Metalens for Virtual Reality Visors

Virtual and augmented reality (VR/AR) visors require compact and lightweight optics. Metalenses have been widely proposed as ultrathin replacements for bulky refractive eyepieces, with performance typically assessed using point spread function (PSF) and modulation transfer function (MTF) measurements. Here, we design, fabricate, and characterize a single-layer silicon nitride metalens optimized for the three emission peaks of an RGB OLED display, and benchmark it against refractive and Fresnel eyepieces. Under coherent illumination, the metalens exhibits a tightly confined PSF and strong mid-to-high spatial-frequency MTF, suggesting excellent optical performance. However, when evaluated in a realistic system-level VR testbed incorporating incoherent OLED illumination, a dynamic-pupil eye model, and near-eye-relevant focal lengths, the same device exhibits pronounced ghosting and background haze. We show that these artifacts arise from the intrinsic multifocal nature of polychromatic diffractive focusing and demonstrate that common mitigation strategies such as narrowband filtering and long-focal-length relay optics merely mask, rather than resolve, the issue. Our results establish that meta-optics for AR/VR must be evaluated under realistic system-level conditions to reveal their true imaging performance.

physics.optics

Improving Human Diving Endurance with a Field-Deployable, Untethered Exoskeleton

Human endurance in underwater locomotion is fundamentally restricted by high energetic demands to overcome drag and the finite supply of self-contained breathing gas. While exoskeleton technology can reduce the metabolic cost of humans in terrestrial locomotion, its potential to enhance human endurance during underwater diving remains entirely unexplored. Here, we present DiveMate, a field-deployable, untethered exoskeleton designed to improve human diving endurance via adaptive kick assistance in real-world underwater environments. During naturalistic diving, DiveMate increases the travel distance using a given energy (breathing gas) by 42.9% and extends dive duration by 54.9% through reducing gas consumption rate. Marked reductions in muscle activation indicate a decrease in physiological exertion, with the net gas consumption rate decreasing by 47.0%. Kinematic characteristics and regularity improvements further underpin efficient energy economy. These results suggest that applying exoskeleton assistance is beneficial for improving human diving endurance and augmenting their ability to explore the aquatic world. This study extends the application frontier of exoskeletons and provides a potential reference for the design and assessment of future underwater assistive devices.

cs.RO

EEGDancer: Dynamic Emotion Latent Space Masked Modeling with Reinforcement Learning for EEG Continuous Emotion Prediction

Continuous electroencephalography (EEG) emotion prediction aims to model the temporal evolution of human emotional states from EEG signals. Unlike conventional discrete emotion recognition, continuous prediction requires capturing long-range temporal dependencies and coherent emotional dynamics. However, existing methods mainly rely on point-wise regression and directly model noisy high-dimensional EEG features, limiting their ability to characterize continuous emotional evolution.To address these challenges, we propose EEGDancer, a dynamic emotional latent space learning framework for continuous EEG emotion prediction. The framework integrates vector-quantized representation learning, masked temporal modeling, and reinforcement learning-based trajectory optimization into a unified architecture.Specifically, a causal spatiotemporal Vector-Quantization Variational Autoencoder (VQ-VAE) is designed to learn structured emotional prototypes and construct a discrete-continuous emotional latent space from EEG signals. Based on the learned latent representations, a Transformer-based masked dynamic modeling strategy captures long-range emotional dependencies and temporal evolution patterns. Furthermore, continuous emotion prediction is formulated as a sequential decision-making problem, and a Soft Actor-Critic (SAC) framework is introduced to optimize emotional prediction trajectories at the sequence level instead of frame-wise local fitting.Extensive experiments on the SEED, SEED-IV, and Long-Term Naturalistic Emotion datasets demonstrate that EEGDancer consistently outperforms existing machine learning and deep learning methods. Ablation studies further verify the effectiveness of the proposed latent space and reinforcement learning-based trajectory optimization for modeling continuous EEG emotional dynamics.

cs.HC

Limits and Trade-Offs of Shift-Invariant Meta-Optical Encoders for Image Compression

Meta-optical encoders can reduce image data before electronic readout or transmission, but engineered point-spread functions (PSFs) do not automatically outperform conventional imaging. We study scene-agnostic, shift-invariant, linear optical encoders using both a fixed total-variation (TV) reconstruction backend and a learned YOLOv8 detection backend. Under a measurement-budget definition of compression ratio that counts all sensed samples across all channels, we compare lens imaging with spatial binning, positive random multi-channel PSFs, signed random kernels, and orthogonal multi-channel kernels. In the low-noise regime, lens-binning gives the highest reconstruction fidelity and strongest YOLOv8 detection metrics at the same compression ratio. Multi-channel encoders, however, degrade more slowly under measurement noise because the measurements are distributed across complementary channels. These results show that, for scene-agnostic incoherent imaging, engineered convolutional PSFs should be justified primarily by robustness, multiplexing, or downstream system constraints, rather than by an expectation that generic wavefront coding will outperform lens-based binning.

physics.optics

Advantages of Broadband Metalenses for Generalizable Image Classification

Optical neural networks (ONNs) are gaining increasing attention to accelerate machine learning tasks. In particular, static meta-optical encoders designed for task-specific pre-processing have demonstrated orders of magnitude smaller energy consumption over purely digital counterparts, albeit at the cost of a slight degradation in classification accuracy. However, a lack of generalizability poses serious challenges for wide deployment of static meta-optical front-ends. Here, we investigate the utility of a single-layer metalens as a meta-optical encoder in ONNs for generalizable image classification. Specifically, we show that a visible-spectrum broadband metalens can achieve image classification accuracy comparable to high-end, sensor-limited optics and consistently outperforms the corresponding hyperboloid baseline across a wide range of sensor pixel sizes and digital backends. We further design an end-to-end optimized single-aperture metasurface for ImageNet classification and observe that the optimization tends to balance the modulation transfer function (MTF) across wavelengths within the sensor-detectable passband. Together, these observations suggest that the preservation of spatial-frequency information is an important factor influencing the performance of ONNs. Our results provide physical insight into the process of task-driven optical optimization and offer practical guidance for the design of high-performance ONNs and meta-optical encoders for generalizable computer vision tasks.

physics.optics

Non-uniform Thermal Conductivity in Nanoscale Multiple Hotspot Systems

Understanding nanoscale hotspot thermal transport is crucial in electronic devices. Contrary to common perception, recent experiments show that closely spaced nanoscale multiple hotspots can enhance heat dissipation. Here, the thermal transport in nanoscale multiple hotspot systems is investigated by solving the phonon Boltzmann transport equation. The local thermal conductivity is proposed to describe the non-uniform spatial distribution of heat transport capability in nanoscale multiple hotspot systems. The maximum value exceeds the uniform heating case by up to 27%, which is attributed to the spatially varying fraction of unscattered phonons emitted from hotspots. Moreover, the effects and mechanisms of hotspot spacing on thermal transport are investigated, showing that reducing the hotspot spacing can enhance the heat flux by up to 40%. This work challenges the conventional view that thermal transport capability is spatially uniform throughout the system and provides fundamental insights for thermal management in high-power-density integrated circuits.

cond-mat.mes-hall

Effect of non-Fourier heat transport on temperature distribution in High Bandwidth Memory

High Bandwidth Memory (HBM), as a key development trend in future memory chip technology, significantly enhances computer performance. At the same time, the thermal challenges arising from its stacked architecture have drawn considerable attention. Most existing studies on HBM thermal management are based on Fourier's law, neglecting the non-Fourier effects introduced by the micro/nanoscale structures within HBM. In this study, the Monte Carlo method (MC) is employed to solve the phonon Boltzmann transport equation (BTE) and investigate the impact of non-Fourier heat transport on the thermal behavior of HBM structures. The results reveal that non-Fourier heat transport leads to a junction temperature that is 59.8 K higher than that predicted by Fourier's law. Furthermore, it is found that the phonon transmittance at the chip interlayers has a severe impact on heat dissipation, with the temperature variation reaching up to 56.6 K. These findings provide more accurate thermal insights, which are critical for the optimized design of HBM systems.

physics.app-ph

Meta-optical Miniscope for Multifunctional Imaging

Miniaturized microscopes (miniscopes) have opened a new frontier in animal behavior studies, enabling real-time imaging of neuron activity while leaving animals largely unconstrained. Canonical designs typically use Gradient-Index (GRIN) lenses or refractive lenses as the objective module for excitation and fluorescence collection, but GRIN lenses suffer from aberrations and refractive lenses are bulky and complex. Meta-optics, composed of subwavelength diffractive elements, offer a promising alternative by combining multiple functionalities with significantly reduced footprint and weight. Here, we present meta-optical miniscopes that integrate functionalities including large field of view (FOV), extended depth of focus (EDOF), and depth sensitivity. These meta-optics replace the traditional refractive lens assembly, reducing the total track length of the objective module from 6.7 mm to 2.5 mm while enhancing imaging performance. Our results demonstrate that meta-optical miniscopes can expand the miniscope toolbox and facilitate the development of more compact and multifunctional imaging systems.

physics.optics

Neural Tangent Knowledge Distillation for Optical Convolutional Networks

Hybrid Optical Neural Networks (ONNs, typically consisting of an optical frontend and a digital backend) offer an energy-efficient alternative to fully digital deep networks for real-time, power-constrained systems. However, their adoption is limited by two main challenges: the accuracy gap compared to large-scale networks during training, and discrepancies between simulated and fabricated systems that further degrade accuracy. While previous work has proposed end-to-end optimizations for specific datasets (e.g., MNIST) and optical systems, these approaches typically lack generalization across tasks and hardware designs. To address these limitations, we propose a task-agnostic and hardware-agnostic pipeline that supports image classification and segmentation across diverse optical systems. To assist optical system design before training, we estimate achievable model accuracy based on user-specified constraints such as physical size and the dataset. For training, we introduce Neural Tangent Knowledge Distillation (NTKD), which aligns optical models with electronic teacher networks, thereby narrowing the accuracy gap. After fabrication, NTKD also guides fine-tuning of the digital backend to compensate for implementation errors. Experiments on multiple datasets (e.g., MNIST, CIFAR, Carvana Masking) and hardware configurations show that our pipeline consistently improves ONN performance and enables practical deployment in both pre-fabrication simulations and physical implementations.

cs.CV

Collaborative On-Sensor Array Cameras

Modern nanofabrication techniques have enabled us to manipulate the wavefront of light with sub-wavelength-scale structures, offering the potential to replace bulky refractive surfaces in conventional optics with ultrathin metasurfaces. In theory, arrays of nanoposts provide unprecedented control over manipulating the wavefront in terms of phase, polarization, and amplitude at the nanometer resolution. A line of recent work successfully investigates flat computational cameras that replace compound lenses with a single metalens or an array of metasurfaces a few millimeters from the sensor. However, due to the inherent wavelength dependence of metalenses, in practice, these cameras do not match their refractive counterparts in image quality for broadband imaging, and may even suffer from hallucinations when relying on generative reconstruction methods. In this work, we investigate a collaborative array of metasurface elements that are jointly learned to perform broadband imaging. To this end, we learn a nanophotonics array with 100-million nanoposts that is end-to-end jointly optimized over the full visible spectrum--a design task that existing inverse design methods or learning approaches cannot support due to memory and compute limitations. We introduce a distributed meta-optics learning method to tackle this challenge. This allows us to optimize a large parameter array along with a learned meta-atom proxy and a non-generative reconstruction method that is parallax-aware and noise-aware. The proposed camera performs favorably in simulation and in all experimental tests irrespective of the scene illumination spectrum.

physics.optics

Accelerating Posterior sampling for Scalable Gaussian Process model

This Paper conducts a thorough simulation study to assess the effectiveness of various acceleration techniques designed to enhance the conjugate gradient algorithm, which is used for solving large linear systems to accelerate Bayesian computation in spatial analysis. The focus is on the application of symbolic decomposition and preconditioners, which are essential for the computational efficiency of conjugate gradient. The findings reveal notable differences in the effectiveness of these acceleration methods. Specific preconditioners, such as the Diagonal Preconditioner, consistently delivered improvements in computational speed. However, in settings involving high-dimensional matrices, traditional solvers were less effective, underscoring the importance of specialized acceleration techniques like the diagonal preconditioner and cgsparse. These methods demonstrated robust performance across a variety of scenarios. The results of this study not only enhance our understanding of the algorithmic dynamics within spatial statistics but also offer valuable guidance for practitioners in choosing the most appropriate computational techniques for their specific needs.

stat.CO

Free Space Few-Photon Nonlinearity in Critically Coupled Polaritonic Metasurfaces

Few-photon optical nonlinearity in planar solid-state systems is challenging yet crucial for quantum and classical optical information processing. Polaritonic nonlinear metasurfaces have emerged as a promising candidate to push the photon number down -- but have often been hindered by challenges like the poor photon-trapping efficiency and lack of modal overlap. Here, we address these issues in a self-hybridized perovskite metasurface through critical coupling engineering, and report strong polaritonic nonlinear absorption at an ultra-low incident power density of only 519 W/cm2 (2 orders of magnitude lower than the state of art in free-space planar devices), with an estimated photon number of 6.12 per cavity lifetime. Taking advantage of a quasi-bound-state-in-the-continuum design with asymmetry-controlled quality-(Q)-factor, we systematically examine the Q-dependent device nonlinearity and determine the optimal cavity critical coupling condition. With the optimized device, we demonstrate at 6 Kelvin a tunable nonlinear response from reverse saturable absorption to saturable absorption at varying pump powers, with a maximal effective nonlinear absorption coefficient up to 29.4+-5.8 cm/W (6 orders of magnitude larger than unpatterned perovskites) at 560 nm wavelength. In addition, the cavity-exciton detuning dependent device response is analyzed and well explained by a phase-space-filling model, elucidating the underlying physics and the origin of giant nonlinearity. Our study paves the way towards practical flat nonlinear optical devices with large functional areas and massive parallel operation capabilities.

physics.optics

An Illumination-Robust Feature Extractor Augmented by Relightable 3D Reconstruction

Visual features, whose description often relies on the local intensity and gradient direction, have found wide applications in robot navigation and localization in recent years. However, the extraction of visual features is usually disturbed by the variation of illumination conditions, making it challenging for real-world applications. Previous works have addressed this issue by establishing datasets with variations in illumination conditions, but can be costly and time-consuming. This paper proposes a design procedure for an illumination-robust feature extractor, where the recently developed relightable 3D reconstruction techniques are adopted for rapid and direct data generation with varying illumination conditions. A self-supervised framework is proposed for extracting features with advantages in repeatability for key points and similarity for descriptors across good and bad illumination conditions. Experiments are conducted to demonstrate the effectiveness of the proposed method for robust feature extraction. Ablation studies also indicate the effectiveness of the self-supervised framework design.

cs.CV

Emotion-Agent: Unsupervised Deep Reinforcement Learning with Distribution-Prototype Reward for Continuous Emotional EEG Analysis

Continuous electroencephalography (EEG) signals are widely used in affective brain-computer interface (aBCI) applications. However, not all continuously collected EEG signals are relevant or meaningful to the task at hand (e.g., wondering thoughts). On the other hand, manually labeling the relevant parts is nearly impossible due to varying engagement patterns across different tasks and individuals. Therefore, effectively and efficiently identifying the important parts from continuous EEG recordings is crucial for downstream BCI tasks, as it directly impacts the accuracy and reliability of the results. In this paper, we propose a novel unsupervised deep reinforcement learning framework, called Emotion-Agent, to automatically identify relevant and informative emotional moments from continuous EEG signals. Specifically, Emotion-Agent involves unsupervised deep reinforcement learning combined with a heuristic algorithm. We first use the heuristic algorithm to perform an initial global search and form prototype representations of the EEG signals, which facilitates the efficient exploration of the signal space and identify potential regions of interest. Then, we design distribution-prototype reward functions to estimate the interactions between samples and prototypes, ensuring that the identified parts are both relevant and representative of the underlying emotional states. Emotion-Agent is trained using Proximal Policy Optimization (PPO) to achieve stable and efficient convergence. Our experiments compare the performance with and without Emotion-Agent. The results demonstrate that selecting relevant and informative emotional parts before inputting them into downstream tasks enhances the accuracy and reliability of aBCI applications.

cs.HC

Wide Field of View Large Aperture Meta-Doublet Eyepiece

Wide field of view and light weight optics are critical for advanced eyewear, with applications in augmented/virtual reality and night vision. Conventional refractive lenses are often stacked to correct aberrations at wide field of view, leading to limited performance and increased size and weight. In particular, simultaneously achieving wide field of view and large aperture for light collection is desirable but challenging to realize in a compact form-factor. Here, we demonstrate a wide field of view (greater than 60$^\circ$) meta-optic doublet eyepiece with an entrance aperture of 2.1 cm. At the design wavelength of 633 nm, the meta-optic doublet achieves comparable performance to a refractive lens-based eyepiece system. This meta-doublet eyepiece illustrates the potential for meta-optics to play an important role in the development of high-quality monochrome near-eye display and night vision systems.

physics.optics

Joint Contrastive Learning with Feature Alignment for Cross-Corpus EEG-based Emotion Recognition

The integration of human emotions into multimedia applications shows great potential for enriching user experiences and enhancing engagement across various digital platforms. Unlike traditional methods such as questionnaires, facial expressions, and voice analysis, brain signals offer a more direct and objective understanding of emotional states. However, in the field of electroencephalography (EEG)-based emotion recognition, previous studies have primarily concentrated on training and testing EEG models within a single dataset, overlooking the variability across different datasets. This oversight leads to significant performance degradation when applying EEG models to cross-corpus scenarios. In this study, we propose a novel Joint Contrastive learning framework with Feature Alignment (JCFA) to address cross-corpus EEG-based emotion recognition. The JCFA model operates in two main stages. In the pre-training stage, a joint domain contrastive learning strategy is introduced to characterize generalizable time-frequency representations of EEG signals, without the use of labeled data. It extracts robust time-based and frequency-based embeddings for each EEG sample, and then aligns them within a shared latent time-frequency space. In the fine-tuning stage, JCFA is refined in conjunction with downstream tasks, where the structural connections among brain electrodes are considered. The model capability could be further enhanced for the application in emotion detection and interpretation. Extensive experimental results on two well-recognized emotional datasets show that the proposed JCFA model achieves state-of-the-art (SOTA) performance, outperforming the second-best method by an average accuracy increase of 4.09% in cross-corpus EEG-based emotion recognition tasks.

cs.HC

HouYi: An open-source large language model specially designed for renewable energy and carbon neutrality field

Renewable energy is important for achieving carbon neutrality goal. With the great success of Large Language Models (LLMs) like ChatGPT in automatic content generation, LLMs are playing an increasingly important role. However, there has not been a specially designed LLM for renewable energy. Meanwhile, there has not been any dataset of renewable energy for training LLMs. Therefore, this paper published the first open-source Renewable Energy Academic Paper (REAP) dataset for non-commercial LLM research of renewable energy. REAP dataset is collected through searching the title and abstract of 1,168,970 academic literatures from Web of Science. Based on REAP dataset, HouYi model, the first LLM for renewable energy, is developed through finetuning general LLMs. HouYi demonstrated powerful academic paper paragraph generation ability in renewable energy field. Experiments show that its ability to generate academic papers on renewable energy is comparable to ChatGPT, slightly outperforms Claude, ERNIE Bot and SparkDesk, and significantly outperforms open-source LLaMA-13B model.

cs.CL

Large Field-of-View Thermal Imaging via All-Silicon Meta-Optics

A broad range of imaging and sensing technologies in the infrared require large Field-of-View (FoV) operation. To achieve this, traditional refractive systems often employ multiple elements to compensate for aberrations, which leads to excess size, weight, and cost. For many applications, including night vision eye-wear, air-borne surveillance, and autonomous navigation for unmanned aerial vehicles, size and weight are highly constrained. Sub-wavelength diffractive optics, also known as meta-optics, can dramatically reduce the size, weight, and cost of these imaging systems, as meta-optics are significantly thinner and lighter than traditional refractive lenses. Here, we demonstrate 80$^\circ$ FoV thermal imaging in the long-wavelength infrared regime (8-12 $\mu$m) using an all-silicon meta-optic with an entrance aperture and lens focal length of 1 cm.

physics.optics