Search arXiv⌕ Search

arXiv subjects

Hyoseok Park

Publications and source records attributed to Hyoseok Park.

6 recordsLinked to original sources

TorchFDTD: GPU-accelerated finite-difference time-domain simulation with discrete adjoints and host-streamed execution for photonic inverse design

Gradient-based photonic design needs full-wave derivatives with respect to millions of parameters, but on a single GPU workstation the time-domain adjoint is limited by device memory and by the incompatibility of fused update kernels with automatic differentiation. Here, we present TorchFDTD, an open-source finite-difference time-domain (FDTD) package that addresses both limits. Its Yee, absorber and dispersion updates execute as fused CUDA kernels captured in a CUDA graph, and every kernel is paired with a transpose kernel derived from its update, so material and geometry derivatives are obtained by a discrete adjoint that PyTorch chains with differentiable objectives built from the supported observations. The adjoint recovers forward states by checkpoint replay or, for lossless periodic problems, by time reversal. For problems that exceed the device, a streamed mode advances the domain one causal slab at a time and keeps the global state and the checkpoints in host memory, which lowers the device allocation while preserving the resident discretization. We validate the package against analytic solutions, against Meep, FDTDX and a rigorous coupled-wave solver, and against automatic differentiation and finite differences. On an A100 the fused path completes full forward solves of eight test scenes 9.0 to 17.1 times faster than the PyTorch FDTD package it extends, and its double-precision solves are 49 to 59 times faster than those of Meep on a workstation CPU. On an RTX 3060, host streaming of $256^3$ and $320^3$ adjoints costs 3.1 and 2.7 times the resident time and lowers the peak device allocation by 56% and 65%. A 54-million-cell pillar-array lens coupled to an angular-spectrum objective yields an adjoint derivative within 0.78% of a central difference. The time-domain adjoint of a device with tens of millions of cells thus becomes available on a single workstation GPU.

physics.optics↗

Imaging-system-aware color routers optimized for imaging information

Conventional nanophotonic color routers are typically optimized under idealized, normal plane waves. However, this standard assumption fails in real-world imaging-system environments, where structures are illuminated by converging light cones and field-dependent chief-ray angles. Here, we present an imaging-system-aware, end-to-end inverse-design framework that directly maximizes the mutual imaging information $\Iimg$ preserved by a single-layer silicon nitride color router under realistic pupil illumination. By analytically embedding the optimal reconstruction decoder directly inside the gradient loop, we co-design the optical nanostructures and the digital recovery pipeline. To scale this approach across a full sensor, we exploit the $D_4$ symmetry of the square pixel lattice, tiling $48$ distinct sensor-field regions using only six unique lithographic masks. Our optimized router is predicted to collect $2.8\times$ more photoelectrons than a conventional color-filter array. Consequently, under low-light conditions, below a green-site signal-to-noise ratio of $13.7$~dB, the color router preserves superior image information compared to the color-filter array; evaluated from its measured routing fractions together with the modeled throughput, the fabricated device reproduces this crossover at $11.1^{+2.1}_{-2.3}$~dB. This marks the first experimental demonstration, from measured routing and a modeled throughput, of a single-layer nanophotonic color router achieving a performance crossover against the color-filter array. These results establish that next-generation flat optics must shift from isolated device efficiency toward system-level co-design optimized under physical imaging-system-pupil geometry.

physics.optics↗

Information-optimized color metalenses for camera imaging

A metalens is conventionally built by prescribing a target phase and matching each meta-atom to it, and making that phase achromatic across the visible band is the main difficulty in color imaging. We remove the prescribed phase entirely. The meta-atom width map is optimized directly against the information the raw, noisy, mosaic-sampled camera measurement carries about the target image, differentiated through Maxwell propagation, the color-filter mosaic, pixel integration and sensor noise, with material, thickness and library held fixed. On a single-layer silicon-nitride platform, starting from a hyperbolic design, this raises delivered target information by 15.8% despite lower collected charge, produces the most balanced set of channel responses at the sensor plane relative to their respective axial maxima, and reduces blue-channel blur. A high-signal-to-noise held-out reconstruction improves in spatial fidelity and color accuracy. The same width-only optimization improves color across silicon-nitride, silicon-dioxide and titanium-dioxide libraries.

physics.optics↗

Coverage-information uncertainty for single-helicity light

A circularly polarized far field must be dark somewhere: its amplitude is a section of a twisted line bundle. We prove this exacts a joint coverage-rate cost that cannot vanish faster than the inverse fourth power of the mode order, and we construct an explicit family of fields attaining that exponent. At lowest order the global optimum is the spin-1 anticoherent state, independently known as optimal for rotations about an unknown axis. The same floor caps photon collection in single-helicity quantum links; the opposite helicity removes it at a cost set by the inverse mode count.

physics.optics↗

Photonic Exponential Approximation via Cascaded TFLN Microring Resonators toward Softmax

The rapid growth of large-scale AI models has intensified energy consumption and data-movement challenges in modern datacenters. Photonic accelerators offer a promising path by executing the linear matrix multiplications of transformer inference at high throughput and low energy. However, the softmax attention layer, which requires element-wise exponentiation followed by normalization, still relies on electronic post-processing, creating an electro-optic conversion bottleneck that negates much of the potential photonic advantage. We present a cascaded micro-ring resonator (MRR) architecture that synthesizes the per-channel exponential function required by softmax, e^{x_n - max(x)}, over a finite interval with tunable worst-case relative error. A control signal detunes each ring via an electro-optic mechanism; a weak probe at fixed frequency experiences Lorentzian transmission, and cascading N identical stages yields a multiplicative transfer function whose logarithm is approximately linear. We derive mapping rules, depth-scaling estimates, and a minimax fitting formulation, and validate the framework with three-dimensional FDTD simulations of X-cut thin-film lithium niobate (TFLN) add-drop micro-ring resonators. Direct multi-ring FDTD validation extends to a five-ring cascade and confirms agreement with theory primarily over the upper operating range; deeper cascades and higher quality factors are assessed analytically. The cascade implements the per-channel exponential block, the key missing nonlinearity for photonic softmax. We further present a WDM-parallel chip architecture with closed-loop PI feedback that completes the full softmax-exponentiation, summation, and normalization-on a single photonic chip without per-channel normalization circuitry.

physics.optics↗

PRISM: Breaking the O(n) Memory Wall in Long-Context LLM Inference via O(1) Photonic Block Selection

Long-context LLM inference is bottlenecked not by compute but by the O(n) memory bandwidth cost of scanning the KV cache at every decode step -- a wall that no amount of arithmetic scaling can break. Recent photonic accelerators have demonstrated impressive throughput for dense attention computation; however, these approaches inherit the same O(n) memory scaling as electronic attention when applied to long contexts. We observe that the real leverage point is the coarse block-selection step: a memory-bound similarity search that determines which KV blocks to fetch. We identify, for the first time, that this task is structurally matched to the photonic broadcast-and-weight paradigm -- the query fans out to all candidates via passive splitting, signatures are quasi-static (matching electro-optic MRR programming), and only rank order matters (relaxing precision to 4-6 bits). Crucially, the photonic advantage grows with context length: as N increases, the electronic scan cost rises linearly while the photonic evaluation remains O(1). We instantiate this insight in PRISM (Photonic Ranking via Inner-product Similarity with Microring weights), a thin-film lithium niobate (TFLN) similarity engine. Hardware-impaired needle-in-a-haystack evaluation on Qwen2.5-7B confirms 100% accuracy from 4K through 64K tokens at k=32, with 16x traffic reduction at 64K context. PRISM achieves a four-order-of-magnitude energy advantage over GPU baselines at practical context lengths (n >= 4K).

physics.optics↗