SearcharxivSearch

arXiv subjects

Hongbo Guo

Publications and source records attributed to Hongbo Guo.

11 recordsLinked to original sources

More Data, Worse Decisions? Preference Reversals in Neural Networks under Gram Incompatibility

Neural networks increasingly combine data across populations, time periods, and operating conditions to improve generalization. This raises a reliability question: whether a model refitted on pooled data preserves an action ordering supported by both sources. Case-Based Decision Theory (CBDT) formalizes this requirement through its composition axiom, which requires source-supported preferences to survive their union. We study when this property holds for fixed-representation neural networks with ordinary least squares (OLS) output heads. First, we show that pooled refitting recomputes the inverse-Gram geometry used to weight source evidence, which can reverse shared preferences, and derive exact and approximate preservation conditions. Next, we introduce a scale-invariant Gram mismatch measure for prioritizing candidate pools and geometry-oriented regularization for shaping source geometry during training. Finally, we develop a three-stage audit that traces strict pairwise reversals through decision changes to task-defined utility loss. Experiments spanning a load-based bidding proxy and medical and financial decision proxies reveal stable and reversal-prone pooling regimes: the load audit identifies a measurable nonzero class of source-consensus-relative harmful decisions under the proxy utility, while cross-domain audits show that comparable mismatch can correspond to sharply different preservation rates. Geometry-oriented objectives occupy distinct descriptive accuracy-consistency-geometry-harm operating points. Together, the framework makes compositional reliability measurable and operational through screening, analytic certification, geometry-oriented training, and decision-consequence auditing.

stat.ML

Mantis: Lightweight Foundation Model for Time Series Classification

While foundation models have revolutionized various domains, their application to time series classification remains rather under-explored, with existing literature predominantly focused on forecasting. To bridge this gap, we introduce \textbf{Mantis}, a transformer-based foundation model pre-trained exclusively on synthetic data via self-supervised contrastive learning. We demonstrate that effective tokenization is critical to unlocking the full potential of transformers, proposing a novel token generator unit. Furthermore, we introduce an enhanced test-time methodology that bridges the performance gap between Mantis and strong specialized approaches by leveraging intermediate-layer representations, self-ensembling, and cross-model embedding fusion. Extensive experiments demonstrate that Mantis establishes a new state-of-the-art, outperforming existing foundation models across four diverse dataset collections covering various application domains.

cs.LG

Pearl: A Production-ready Reinforcement Learning Agent

Reinforcement learning (RL) is a versatile framework for optimizing long-term goals. Although many real-world problems can be formalized with RL, learning and deploying a performant RL policy requires a system designed to address several important challenges, including the exploration-exploitation dilemma, partial observability, dynamic action spaces, and safety concerns. While the importance of these challenges has been well recognized, existing open-source RL libraries do not explicitly address them. This paper introduces Pearl, a Production-Ready RL software package designed to embrace these challenges in a modular way. In addition to presenting benchmarking results, we also highlight examples of Pearl's ongoing industry adoption to demonstrate its advantages for production use cases. Pearl is open sourced on GitHub at github.com/facebookresearch/pearl and its official website is pearlagent.github.io.

cs.LG

LOGCAN++: Adaptive Local-global class-aware network for semantic segmentation of remote sensing imagery

Remote sensing images usually characterized by complex backgrounds, scale and orientation variations, and large intra-class variance. General semantic segmentation methods usually fail to fully investigate the above issues, and thus their performances on remote sensing image segmentation are limited. In this paper, we propose our LOGCAN++, a semantic segmentation model customized for remote sensing images, which is made up of a Global Class Awareness (GCA) module and several Local Class Awareness (LCA) modules. The GCA module captures global representations for class-level context modeling to reduce the interference of background noise. The LCA module generates local class representations as intermediate perceptual elements to indirectly associate pixels with the global class representations, targeting at dealing with the large intra-class variance problem. In particular, we introduce affine transformations in the LCA module for adaptive extraction of local class representations to effectively tolerate scale and orientation variations in remotely sensed images. Extensive experiments on three benchmark datasets show that our LOGCAN++ outperforms current mainstream general and remote sensing semantic segmentation methods and achieves a better trade-off between speed and accuracy. Code is available at https://github.com/xwmaxwma/rssegmentation.

cs.CV

Uncertainty of Joint Neural Contextual Bandit

Contextual bandit learning is increasingly favored in modern large-scale recommendation systems. To better utlize the contextual information and available user or item features, the integration of neural networks have been introduced to enhance contextual bandit learning and has triggered significant interest from both academia and industry. However, a major challenge arises when implementing a disjoint neural contextual bandit solution in large-scale recommendation systems, where each item or user may correspond to a separate bandit arm. The huge number of items to recommend poses a significant hurdle for real world production deployment. This paper focuses on a joint neural contextual bandit solution which serves all recommending items in one single model. The output consists of a predicted reward $μ$, an uncertainty $σ$ and a hyper-parameter $α$ which balances exploitation and exploration, e.g., $μ+ ασ$. The tuning of the parameter $α$ is typically heuristic and complex in practice due to its stochastic nature. To address this challenge, we provide both theoretical analysis and experimental findings regarding the uncertainty $σ$ of the joint neural contextual bandit model. Our analysis reveals that $α$ demonstrates an approximate square root relationship with the size of the last hidden layer $F$ and inverse square root relationship with the amount of training data $N$, i.e., $σ\propto \sqrt{\frac{F}{N}}$. The experiments, conducted with real industrial data, align with the theoretical analysis, help understanding model behaviors and assist the hyper-parameter tuning during both offline training and online deployment.

cs.LG

Multi-Scale and Multi-Modal Contrastive Learning Network for Biomedical Time Series

Multi-modal biomedical time series (MBTS) data offers a holistic view of the physiological state, holding significant importance in various bio-medical applications. Owing to inherent noise and distribution gaps across different modalities, MBTS can be complex to model. Various deep learning models have been developed to learn representations of MBTS but still fall short in robustness due to the ignorance of modal-to-modal variations. This paper presents a multi-scale and multi-modal biomedical time series representation learning (MBSL) network with contrastive learning to migrate these variations. Firstly, MBTS is grouped based on inter-modal distances, then each group with minimum intra-modal variations can be effectively modeled by individual encoders. Besides, to enhance the multi-scale feature extraction (encoder), various patch lengths and mask ratios are designed to generate tokens with semantic information at different scales and diverse contextual perspectives respectively. Finally, cross-modal contrastive learning is proposed to maximize consistency among inter-modal groups, maintaining useful information and eliminating noises. Experiments against four bio-medical applications show that MBSL outperforms state-of-the-art models by 33.9% mean average errors (MAE) in respiration rate, by 13.8% MAE in exercise heart rate, by 1.41% accuracy in human activity recognition, and by 1.14% F1-score in obstructive sleep apnea-hypopnea syndrome.

cs.LG

Evaluating Online Bandit Exploration In Large-Scale Recommender System

Bandit learning has been an increasingly popular design choice for recommender system. Despite the strong interest in bandit learning from the community, there remains multiple bottlenecks that prevent many bandit learning approaches from productionalization. One major bottleneck is how to test the effectiveness of bandit algorithm with fairness and without data leakage. Different from supervised learning algorithms, bandit learning algorithms emphasize greatly on the data collection process through their explorative nature. Such explorative behavior may induce unfair evaluation in a classic A/B test setting. In this work, we apply upper confidence bound (UCB) to our large scale short video recommender system and present a test framework for the production bandit learning life-cycle with a new set of metrics. Extensive experiment results show that our experiment design is able to fairly evaluate the performance of bandit learning in the recommender system.

cs.IR

Autoignition of two-phase n-heptane/air mixtures behind an oblique shock: insights into spray oblique detonation initiation

Autoignition of n-heptane droplet/vapor/air mixtures behind an oblique shock wave are studied, through Eulerian-Lagrangian method and a skeletal chemical mechanism. The effects of gas/liquid equivalence ratio (ER), droplet diameter, flight altitude, and Mach number on the ignition transient and chemical timescales are investigated. The results show that the ratio of chemical excitation time to ignition delay time can be used to predict the oblique detonation wave (ODW) transition mode. When the ratio is relatively high, the combustion heat release is slow and smooth transition is more likely to occur. In heterogeneous ignition, there are direct interactions between the evaporating droplets and the induction/ignition process, and the chemical explosive propensity changes accordingly. The energy absorption of evaporating droplets significantly retards the ignition of n-heptane vapor. In the two-phase n-heptane mixture autoignition process, the ignition delay time decreases exponentially with flight Mach number, and increases first and then decreases with the flight altitude. As the liquid ER increases, both ignition delay time and droplet evaporation time increase. With increased droplet diameter, the ignition delay time decreases, and the evaporation time increases. Besides, for Mach number is less than 10, the ratio of the chemical excitation time to ignition delay time generally increases with the flight altitude or Mach number. It increases when the liquid ER decreases or droplet diameter increases. When Mach number is sufficiently high, it shows limited change with fuel and inflow conditions. The results from this work can provide insights into spray ODW initiation. The ODW is more likely to be initiated with a smooth transition at high altitude or Mach number. Abrupt transition mode tends to happen when fine fuel droplets are loaded.

physics.flu-dyn

Ignition limit and shock-to-detonation transition mode of n-heptane/air mixture in high-speed wedge flows

In this work, oblique detonation of n-heptane/air mixture in high-speed wedge flows is simulated by solving the reactive Euler equations with a two-dimensional (2D) configuration. This is a first attempt to model complicated hydrocarbon fuel ODWs with a detailed chemistry (44 species and 112 reactions). Effects of freestream equivalence ratios and velocities are considered, and the abrupt and smooth transition from oblique shock to detonation are predicted. Ignition limit, ODW characteristics, and predictability of the transition mode are discussed. Firstly, homogeneous constant-volume ignition calculations are performed for both fuel-lean and stoichiometric mixtures. The results show that the ignition delay generally increases with the wedge angle. However, a negative wedge angle dependence is observed, due to the negative temperature coefficient effects. The wedge angle range for successful ignition of n-heptane/air mixtures decreases when the wedge length is reduced. From 2D simulations of stationary ODWs, the initiation length generally decreases with the freestream equivalence ratio, but the transition length exhibits weakly non-monotonic dependence. Smooth ODW typically occurs for lean conditions (equivalence ratio < 0.4). The interactions between shock / compression waves and chemical reaction inside the induction zone are also studied with the chemical explosive mode analysis. Moreover, the predictability of the shock-to-detonation transition mode is explored through quantifying the relation between ignition delay and chemical excitation time. It is demonstrated that the ignition delay (excitation time) increases (decreases) with the freestream equivalence ratio for the three studied oncoming flow velocities. Smaller excitation time corresponds to stronger pressure waves from the ignition location behind OSW.

physics.flu-dyn

On the evolutions of induction zone structure in wedge-stabilized oblique detonation with water mist flows

Two-dimensional wedge-stabilized oblique detonations in stoichiometric and fuel-lean H2/O2/Ar mixtures with water mists are studied with Eulerian-Lagrangian method. The effects of water droplet mass flow rate on flow and chemical structures in the induction zone, as well as physical / chemical roles of water vapor, are investigated. The results show that the oblique detonation wave (ODW) can stand in a range of water mass flow rates for both stoichiometric and fuel-lean mixtures. With increased droplet mass flow rate, the deflagration front in the induction zone is distorted and becomes zigzagged, but the transition mode from oblique shock wave (OSW) to ODW does not change. Moreover, the initiation and transition locations monotonically increase, and the OSW and ODW angles decrease, due to droplet evaporation and water vapor dilution in the induction region. For fuel-lean mixtures, the sensitivity of characteristic locations to the droplet loading variations is mild, which signifies better intrinsic stability and resilience to the oncoming water droplets. The chemical explosiveness of the gaseous mixture between the lead shock and reaction front is studied with the chemical explosive method analysis. The smooth transition is caused by the highly enhanced reactivity of the gas immediately behind the curved shock, intensified by the compression waves. Nonetheless, the abrupt transition results from the intersection between the beforehand generated detonation wave in the induction zone and OSW. Besides, the degree to which the gas chemical reactivity in the induction zone for fuel-lean mixtures is reduced by evaporating droplets is generally lower than that for stoichiometric gas. Also, physical and chemical effects of water vapor from liquid droplets result in significant differences in ODW initiation and morphology.

physics.flu-dyn

Polynomial maps with invertible sums of Jacobian matrices and of directional Derivatives

Let $F: C^n \rightarrow C^m$ be a polynomial map with $degF=d \geq 2$. We prove that $F$ is invertible if $m = n$ and $\sum^{d-1}_{i=1} JF(α_i)$ is invertible for all $i$, which is trivially the case for invertible quadratic maps. More generally, we prove that for affine lines $L = \{β+ μγ| μ\in C\} \subseteq C^n$ ($γ\ne 0$), $F|_L$ is linearly rectifiable, if and only if $\sum^{d-1}_{i=1} JF(α_i) \cdot γ\ne 0$ for all $α_i \in L$. This appears to be the case for all affine lines $L$ when $F$ is injective and $d \le 3$. We also prove that if $m = n$ and $\sum^{n}_{i=1} JF(α_i)$ is invertible for all $α_i \in C^n$, then $F$ is a composition of an invertible linear map and an invertible polynomial map $X+H$ with linear part $X$, such that the subspace generated by $\{JH(α) | α\in C^n\}$ consists of nilpotent matrices.

math.AC