SearcharxivSearch

arXiv subjects

Michael Huang

Publications and source records attributed to Michael Huang.

At least 19 recordsLinked to original sources

Integrated Hardware Annealing based on Langevin Dynamics for Ising Machines

Ising machines are non-von Neumann machines designed to solve combinatorial optimization problems (COP) by searching for the ground state, or the lowest energy configuration, within the Ising model. However, Ising machines often face the challenges of getting trapped in local minima due to the complex energy landscapes. Hardware annealing algorithms help mitigate this issue by using a probabilistic approach to steer the system toward the ground state. In this paper, we present a hardware annealing algorithm for Ising machines based on Langevin dynamics, a stochastic perturbation by random noise. Theoretical analysis, system-level design, and detailed circuit design are carried out. We evaluate the performance of the algorithm through chip-level simulation using a standard 65-nm CMOS technology to demonstrate the algorithm's efficacy. The results show that the proposed hardware annealing algorithm effectively guides the system to reach the ground state with a probability of 86.5%, significantly improving the solution quality by 97.5%. Further, we compare the algorithm with state-of-the-art hardware annealing methods through behavioral-level simulations, highlighting its improved solution quality alongside a 50% reduction in time-to-solution.

cs.AR

Conditional Dynamical Systems for Image Generation

Image generation has been dominated by deep generative models running on GPUs, a paradigm whose computational and energy costs raise growing sustainability concerns. Emerging non-von Neumann computing substrates, including quantum, compute-in-memory, photonic, and thermodynamic platforms, promise greater efficiency, yet much of the existing work ports conventional neural architectures onto them and primarily accelerates operations such as matrix multiplication. This does not fully exploit a native capability of many emerging computing substrates: relaxation toward low-energy states can itself perform computation at negligible cost. We develop a family of continuous dynamical systems for image generation, built around this primitive to better harness its computational power. The proposed generator evolves an internal state under dynamics admitting an explicit Lyapunov energy and then renders the resulting state through a compact, class-agnostic decoder. For conditional generation, we introduce energy tilting: programmed pairwise interactions remain fixed and shared across classes, while a class-dependent linear field reshapes the energy without reprogramming the interaction array. An Ising-inspired design reaches a clean-FID of 9.71 on CIFAR-10 with 4096 spin variables. These results suggest that the energy-descending dynamics can serve directly as a generative computation and offer a promising path toward efficient generative tasks beyond GPUs.

cs.ET

Convex hypersurfaces and robust heterodimensional dynamics

We prove that any closed orientable hypersurface in a contact manifold of dimension five or greater is isotopic to a robustly non-convex hypersurface via an arbitrarily $C^0$-small isotopy. This strengthens a recent result of the first author and yields a strong counterpart to the groundbreaking density theorem of Honda-Huang and Giroux. This is proven by combining a new convexity obstruction via heteroclinics and recent advances in robust heterodimensional dynamics due to Li-Turaev to produce a robust deconvexifying plug, which is a local and robust convexity obstruction.

math.SG

Kimodo: Scaling Controllable Human Motion Generation

High-quality human motion data is becoming increasingly important for applications in robotics, simulation, and entertainment. Recent generative models offer a potential data source, enabling human motion synthesis through intuitive inputs like text prompts or kinematic constraints on poses. However, the small scale of public mocap datasets has limited the motion quality, control accuracy, and generalization of these models. In this work, we introduce Kimodo, an expressive and controllable kinematic motion diffusion model trained on 700 hours of optical motion capture data. Our model generates high-quality motions while being easily controlled through text and a comprehensive suite of kinematic constraints including full-body keyframes, sparse joint positions/rotations, 2D waypoints, and dense 2D paths. This is enabled through a carefully designed motion representation and two-stage denoiser architecture that decomposes root and body prediction to minimize motion artifacts while allowing for flexible constraint conditioning. Experiments on the large-scale mocap dataset justify key design decisions and analyze how the scaling of dataset size and model size affect performance.

cs.CV

Nonlinear Dynamical Friction from the Doppler-Shifted Equilibrium Memory Kernel

We present a statistical mechanics framework for modeling equilibrium friction coefficients using the Generalized Langevin Equation (GLE). We show that the kernel, obtained via the Fluctuation-Dissipation Theorem (FDT) from the stochastic force autocorrelation measured in a thermal equilibrium state, is sufficient to model the dynamics of the system in a Non-Equilibrium Steady State (NESS). This approach provides a computationally efficient path to modeling complex equilibrium friction problems. We apply this framework to the canonical problem of test particle drag in a uniform plasma. The GLE formalism is shown to naturally capture non-Markovian phenomena through the moments of the kernel, including an effective mass renormalization and oscillatory relaxation. We demonstrate that the standard Chandrasekhar stopping power formula arises naturally as the Markovian limit of this equilibrium memory kernel. These theoretical predictions are quantitatively validated by direct Particle-in-Cell simulations, which confirm the predicted oscillatory structure of the memory kernel. This work thus establishes a practical method for predicting equilibrium friction properties from first-principles equilibrium simulations.

physics.plasm-ph

Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars

Audio-driven facial animation presents an effective solution for animating digital avatars. In this paper, we detail the technical aspects of NVIDIA Audio2Face-3D, including data acquisition, network architecture, retargeting methodology, evaluation metrics, and use cases. Audio2Face-3D system enables real-time interaction between human users and interactive avatars, facilitating facial animation authoring for game characters. To assist digital avatar creators and game developers in generating realistic facial animations, we have open-sourced Audio2Face-3D networks, SDK, training framework, and example dataset.

cs.GR

Bringing Multimodality to Amazon Visual Search System

Image to image matching has been well studied in the computer vision community. Previous studies mainly focus on training a deep metric learning model matching visual patterns between the query image and gallery images. In this study, we show that pure image-to-image matching suffers from false positives caused by matching to local visual patterns. To alleviate this issue, we propose to leverage recent advances in vision-language pretraining research. Specifically, we introduce additional image-text alignment losses into deep metric learning, which serve as constraints to the image-to-image matching loss. With additional alignments between the text (e.g., product title) and image pairs, the model can learn concepts from both modalities explicitly, which avoids matching low-level visual features. We progressively develop two variants, a 3-tower and a 4-tower model, where the latter takes one more short text query input. Through extensive experiments, we show that this change leads to a substantial improvement to the image to image matching problem. We further leveraged this model for multimodal search, which takes both image and reformulation text queries to improve search quality. Both offline and online experiments show strong improvements on the main metrics. Specifically, we see 4.95% relative improvement on image matching click through rate with the 3-tower model and 1.13% further improvement from the 4-tower model.

cs.CV

Diff-PIC: Revolutionizing Particle-In-Cell Nuclear Fusion Simulation with Diffusion Models

The rapid development of AI highlights the pressing need for sustainable energy, a critical global challenge for decades. Nuclear fusion, generally seen as an ultimate solution, has been the focus of intensive research for nearly a century, with investments reaching hundreds of billions of dollars. Recent advancements in Inertial Confinement Fusion have drawn significant attention to fusion research, in which Laser-Plasma Interaction (LPI) is critical for ensuring fusion stability and efficiency. However, the complexity of LPI upon fusion ignition makes analytical approaches impractical, leaving researchers depending on extremely computation-demanding Particle-in-Cell (PIC) simulations to generate data, presenting a significant bottleneck to advancing fusion research. In response, this work introduces Diff-PIC, a novel framework that leverages conditional diffusion models as a computationally efficient alternative to PIC simulations for generating high-fidelity scientific LPI data. In this work, physical patterns captured by PIC simulations are distilled into diffusion models associated with two tailored enhancements: (1) To effectively capture the complex relationships between physical parameters and corresponding outcomes, the parameters are encoded in a physically-informed manner. (2) To further enhance efficiency while maintaining high fidelity and physical validity, the rectified flow technique is employed to transform our model into a one-step conditional diffusion model. Experimental results show that Diff-PIC achieves 16,200$\times$ speedup compared to traditional PIC on a 100 picosecond simulation, with an average reduction in MAE / RMSE / FID of 59.21% / 57.15% / 39.46% with respect to two other SOTA data generation approaches.

physics.comp-ph

Inertial Confinement Fusion Forecasting via Large Language Models

Controlled fusion energy is deemed pivotal for the advancement of human civilization. In this study, we introduce $\textbf{LPI-LLM}$, a novel integration of Large Language Models (LLMs) with classical reservoir computing paradigms tailored to address a critical challenge, Laser-Plasma Instabilities ($\texttt{LPI}$), in Inertial Confinement Fusion ($\texttt{ICF}$). Our approach offers several key contributions: Firstly, we propose the $\textit{LLM-anchored Reservoir}$, augmented with a $\textit{Fusion-specific Prompt}$, enabling accurate forecasting of $\texttt{LPI}$-generated-hot electron dynamics during implosion. Secondly, we develop $\textit{Signal-Digesting Channels}$ to temporally and spatially describe the driver laser intensity across time, capturing the unique characteristics of $\texttt{ICF}$ inputs. Lastly, we design the $\textit{Confidence Scanner}$ to quantify the confidence level in forecasting, providing valuable insights for domain experts to design the $\texttt{ICF}$ process. Extensive experiments demonstrate the superior performance of our method, achieving 1.90 CAE, 0.14 $\texttt{top-1}$ MAE, and 0.11 $\texttt{top-5}$ MAE in predicting Hard X-ray ($\texttt{HXR}$) energies emitted by the hot electrons in $\texttt{ICF}$ implosions, which presents state-of-the-art comparisons against concurrent best systems. Additionally, we present $\textbf{LPI4AI}$, the first $\texttt{LPI}$ benchmark based on physical experiments, aimed at fostering novel ideas in $\texttt{LPI}$ research and enhancing the utility of LLMs in scientific exploration. Overall, our work strives to forge an innovative synergy between AI and $\texttt{ICF}$ for advancing fusion energy.

cs.LG

Leveraging Large Language Models for Multimodal Search

Multimodal search has become increasingly important in providing users with a natural and effective way to ex-press their search intentions. Images offer fine-grained details of the desired products, while text allows for easily incorporating search modifications. However, some existing multimodal search systems are unreliable and fail to address simple queries. The problem becomes harder with the large variability of natural language text queries, which may contain ambiguous, implicit, and irrelevant in-formation. Addressing these issues may require systems with enhanced matching capabilities, reasoning abilities, and context-aware query parsing and rewriting. This paper introduces a novel multimodal search model that achieves a new performance milestone on the Fashion200K dataset. Additionally, we propose a novel search interface integrating Large Language Models (LLMs) to facilitate natural language interaction. This interface routes queries to search systems while conversationally engaging with users and considering previous searches. When coupled with our multimodal search model, it heralds a new era of shopping assistants capable of offering human-like interaction and enhancing the overall search experience.

cs.CV

Decision-Focused Learning with Directional Gradients

We propose a novel family of decision-aware surrogate losses, called Perturbation Gradient (PG) losses, for the predict-then-optimize framework. The key idea is to connect the expected downstream decision loss with the directional derivative of a particular plug-in objective, and then approximate this derivative using zeroth order gradient techniques. Unlike the original decision loss which is typically piecewise constant and discontinuous, our new PG losses is a Lipschitz continuous, difference of concave functions that can be optimized using off-the-shelf gradient-based methods. Most importantly, unlike existing surrogate losses, the approximation error of our PG losses vanishes as the number of samples grows. Hence, optimizing our surrogate loss yields a best-in-class policy asymptotically, even in misspecified settings. This is the first such result in misspecified settings, and we provide numerical evidence confirming our PG losses substantively outperform existing proposals when the underlying model is misspecified.

cs.LG

Efficient LDPC Decoding using Physical Computation

Due to 5G deployment, there is significant interest in LDPC decoding. While much research is devoted on efficient hardwiring of algorithms based on Belief Propagation (BP), it has been shown that LDPC decoding can be formulated as a combinatorial optimization problem, which could benefit from significant acceleration of physical computation mechanisms such as Ising machines. This approach has so far resulted in poor performance. This paper shows that the reason is not fundamental but suboptimal hardware and formulation. A co-designed Ising machine-based system can improve speed by 3 orders of magnitude. As a result, a physical computation approach can outperform hardwiring state-of-the-art algorithms. In this paper, we show such an augmented Ising machine that is 4.4$\times$ more energy efficient than the state of the art in the literature.

cs.IT

Positive 2-bridge knots and chirally cosmetic surgeries

In this paper we verify that with the exception of the $(2, 2n+1)$ torus knots, positive 2-bridge knots up to 31 crossings do not admit chirally cosmetic surgeries. A knot $K$ admits chirally cosmetic surgeries if there exist surgeries $S^3_r$ and $S^3_{r'}$ with distinct slopes $r$ and $r'$ such that $S^3_r(K) \cong -S^3_{r'}(K)$, where the negative represents an orientation reversal. To verify this, we use the obstruction formula from arXiv:2112.03144 which relates classical knot invariants to the existence of chirally cosmetic surgeries. To check the formula, we develop a Python program that computes the classical knot invariants $a_2$, $a_4$, $v_3$, $\det$, and $g$ of a positive 2-bridge knot.

math.GT

Augmented Electronic Ising Machine as an Effective SAT Solver

With the slowdown of improvement in conventional von Neumann systems, increasing attention is paid to novel paradigms such as Ising machines. They have very different approach to NP-complete optimization problems. Ising machines have shown great potential in solving binary optimization problems like MaxCut. In this paper, we present an analysis of these systems in satisfiability (SAT) problems. We demonstrate that, in the case of 3-SAT, a basic architecture fails to produce meaningful acceleration, thanks in no small part to the relentless progress made in conventional SAT solvers. Nevertheless, careful analysis attributes part of the failure to the lack of two important components: cubic interactions and efficient randomization heuristics. To overcome these limitations, we add proper architectural support for cubic interaction on a state-of-the-art Ising machine. More importantly, we propose a novel semantic-aware annealing schedule that makes the search-space navigation much more efficient than existing annealing heuristics. With experimental analyses, we show that such an Augmented Ising Machine for SAT (AIMS), outperforms state-of-the-art software-based, GPU-based and conventional hardware SAT solvers by orders of magnitude. We also demonstrate AIMS to be relatively robust against device variation and noise.

cs.AI

Supporting Energy-Based Learning With An Ising Machine Substrate: A Case Study on RBM

Nature apparently does a lot of computation constantly. If we can harness some of that computation at an appropriate level, we can potentially perform certain type of computation (much) faster and more efficiently than we can do with a von Neumann computer. Indeed, many powerful algorithms are inspired by nature and are thus prime candidates for nature-based computation. One particular branch of this effort that has seen some recent rapid advances is Ising machines. Some Ising machines are already showing better performance and energy efficiency for optimization problems. Through design iterations and co-evolution between hardware and algorithm, we expect more benefits from nature-based computing systems. In this paper, we make a case for an augmented Ising machine suitable for both training and inference using an energy-based machine learning algorithm. We show that with a small change, the Ising substrate accelerate key parts of the algorithm and achieve non-trivial speedup and efficiency gain. With a more substantial change, we can turn the machine into a self-sufficient gradient follower to virtually complete training entirely in hardware. This can bring about 29x speedup and about 1000x reduction in energy compared to a Tensor Processing Unit (TPU) host.

cs.ET

Debiasing In-Sample Policy Performance for Small-Data, Large-Scale Optimization

Motivated by the poor performance of cross-validation in settings where data are scarce, we propose a novel estimator of the out-of-sample performance of a policy in data-driven optimization.Our approach exploits the optimization problem's sensitivity analysis to estimate the gradient of the optimal objective value with respect to the amount of noise in the data and uses the estimated gradient to debias the policy's in-sample performance. Unlike cross-validation techniques, our approach avoids sacrificing data for a test set, utilizes all data when training and, hence, is well-suited to settings where data are scarce. We prove bounds on the bias and variance of our estimator for optimization problems with uncertain linear objectives but known, potentially non-convex, feasible regions. For more specialized optimization problems where the feasible region is "weakly-coupled" in a certain sense, we prove stronger results. Specifically, we provide explicit high-probability bounds on the error of our estimator that hold uniformly over a policy class and depends on the problem's dimension and policy class's complexity. Our bounds show that under mild conditions, the error of our estimator vanishes as the dimension of the optimization problem grows, even if the amount of available data remains small and constant. Said differently, we prove our estimator performs well in the small-data, large-scale regime. Finally, we numerically compare our proposed method to state-of-the-art approaches through a case-study on dispatching emergency medical response services using real data. Our method provides more accurate estimates of out-of-sample performance and learns better-performing policies.

math.OC

A CMOS-compatible Ising Machine with Bistable Nodes

Physical Ising machines rely on nature to guide a dynamical system towards an optimal state which can be read out as a heuristical solution to a combinatorial optimization problem. Such designs that use nature as a computing mechanism can lead to higher performance and/or lower operation costs and hence have attracted research and prototyping efforts from industry and academia. Quantum annealers are a prominent example of such efforts. However, some physics-centric Ising machines require stringent operating conditions that result in significant bulk and energy budget. Such disadvantages may be acceptable if these designs provide some significant intrinsic advantages at a much larger scale in the future, which remains to be seen. But for now, integrated electronic designs of Ising machines allow more immediate applications. We propose one such design that uses bistable nodes, coupled with programmable and variable strengths. The design is fully CMOS compatible for chip-scale applications and demonstrates competitive solution quality and significantly superior execution time and energy.

quant-ph

Original Research By Young Twinkle Students (ORBYTS): Ephemeris Refinement of Transiting Exoplanets

We report follow-up observations of transiting exoplanets that have either large uncertainties (>10 minutes) in their transit times or have not been observed for over three years. A fully robotic ground-based telescope network, observations from citizen astronomers and data from TESS have been used to study eight planets, refining their ephemeris and orbital data. Such follow-up observations are key for ensuring accurate transit times for upcoming ground and space-based telescopes which may seek to characterise the atmospheres of these planets. We find deviations from the expected transit time for all planets, with transits occurring outside the 1 sigma uncertainties for seven planets. Using the newly acquired observations, we subsequently refine their periods and reduce the current predicted ephemeris uncertainties to 0.28 - 4.01 minutes. A significant portion of this work has been completed by students at two high schools in London as part of the Original Research By Young Twinkle Students (ORBYTS) programme.

astro-ph.EP