SearcharxivSearch

arXiv subjects

Pengyu Liu

Publications and source records attributed to Pengyu Liu.

At least 19 recordsLinked to original sources

The JWST Early Release Science Program for Direct Observations of Exoplanetary Systems VIII: patchy forsterite and enstatite clouds in the atmosphere of VHS 1256 b, retrieval lessons learned and outlook to the future

JWST defines a new era for the data-driven approach of retrieval modelling, which has become a cornerstone tool for the statistical inference of exoplanetary and brown dwarf properties. The Early Release Science program #1386 observations of VHS 1256 b represent a huge jump in data quality, data quantity and spectral coverage for such objects. VHS 1256 b is a young, planetary mass and extremely variable companion that populates the enigmatic L/T cohort of substellar atmospheres. In this first retrieval analysis of the full 1 - 18 micron dataset, we apply the Brewster retrieval framework to the NIRSpec and MIRI spectroscopic observations of VHS 1256 b, exploring a variety of cloud species and structures. Using Delta(BIC) we find that the data is best described by a forsterite (Mg$_{2}$SiO$_{4}$) and enstatite (MgSiO$_{3}$) cloud combination. Our analysis shows a strong preference for patchy silicate cloud coverage, which aligns with VHS 1256 b's extensive and well documented spectral variability. Our retrieval is able to place constraints on the abundances of H$_{2}$O, CO, CO$_{2}$, CH$_{4}$ as well as NH$_{3}$. We also show that the retrieved parameters are sensitive to the data used and the relative signal-to-noise ratios between data from different instruments. We conclude with the next steps for the wider retrieval community to better understand young and cloudy exoplanetary atmospheres.

astro-ph.EP

MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding

Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communication. However, their short duration, weak motion patterns, and fine-grained semantic differences make them difficult to annotate, model, and evaluate in a standardized manner. To promote academic research on micro-action analysis, we proposed and have annually organized the Micro-Action Analysis Grand Challenge (MAC) as a public benchmark platform for this emerging field. The first two editions of MAC established standardized evaluation settings for micro-action recognition and detection, providing publicly accessible datasets and protocols. Building upon these editions, this paper presents the 3rd MAC, held in conjunction with ACM Multimedia 2026. Under the theme of moving from recognition to fine-grained micro-action understanding, this edition further expands the scope of the challenge beyond conventional recognition and detection. In particular, we introduce a new task named fine-grained micro-action understanding, evaluated with the assistance of multimodal large language models, aiming to assess models' ability to capture fine-grained semantic cues and interpret subtle human micro-actions at a deeper level. We summarize the datasets, task settings, evaluation protocols, competition results, and representative solutions from top-performing teams. Finally, we discuss future directions for micro-action analysis and its broader role in human-centric video understanding.

cs.CV

Photometric Variability and Rotation of Beta Pictoris b from JWST NIRCam Coronagraphic Imaging

We report the detection of photometric variability in the directly imaged super-Jupiter $β$ Pictoris b. Using JWST NIRCam dual-band coronagraphic imaging, we conducted a 16-hour continuous photometric monitoring campaign in the F210M and F410M filters. We developed and validated a time-series photometry framework that combines PSF subtraction, principal component analysis for systematic noise removal, and injection-and-recovery tests to confirm signal fidelity. Both light curves show consistent sinusoidal variability at $\sim$5$σ$ and $\gg 5σ$ significance in the F210M and F410M bands, respectively. A joint sinusoidal fit yields a rotation period of $P_{\rm rot} = 9.00 \pm 0.13$ hr and variability amplitudes of $0.85 \pm 0.07\%$ and $0.89 \pm 0.04\%$ in F210M and F410M, respectively. The near-identical amplitudes and periods in both bands confirm a common astrophysical origin in a heterogeneous atmosphere. Combining $P_{\rm rot}$ with the previously measured projected rotational velocity, we constrain the line-of-sight spin axis inclination of $β$ Pic b. The result favors an equator-on viewing geometry, consistent with line-of-sight spin-orbit alignment: the planetary spin axis, orbital plane, debris disk, and stellar equator are all mutually aligned. This stands in sharp contrast to the large obliquities of wide-separation companions that are likely formed via gravitational fragmentation. Together with the system's young age, this observation provides independent dynamical evidence that $β$ Pic b formed via core accretion. This result constitutes the first detection of rotational modulation in a close-in, high-contrast exoplanet that likely formed via core accretion, demonstrating that time-series coronagraphic imaging with JWST opens a powerful new window onto the rotation, atmospheric dynamics, and spin-orbit architecture of this population.

astro-ph.EP

Direct Imaging Discovery of Giant Exoplanet $β$ Pictoris d: A Decade-Long Game of Hide-and-Seek

We report the direct imaging discovery of a third exoplanet in the $β$ Pictoris system. We detect $β$ Pictoris d ($β$ Pic d) in non-coronagraphic observations obtained with VLT/ERIS as well as multi-epoch archival datasets from JWST/NIRCam and VLT/SPHERE. Astrometric measurements over an 11-year baseline demonstrate that it is consistent with a gravitationally-bound source with orbital motion. Joint multi-planet orbit fits of all three planets in the system yield a semi-major axis of $26.0^{+2.2}_{-6.1}$ au and inclination $89.0^{+0.7}_{-0.6}$ deg for planet d. $β$ Pic d has a larger orbital semi-major axis than the other known planets in the system, but is coplanar with the inner two planets, and its orbit is consistent with sculpting the inner edge of the debris disk. $β$ Pic d has a contrast of $ΔL^{\prime}=12.11\pm0.15$ mag, with colors and luminosity that closely match those of 51 Eri b, another exoplanet in the $β$ Pictoris moving group. Its VLT/ERIS and JWST/NIRCam colors are distinct from those of free-floating planetary-mass objects of a similar age and temperature. Its red $F410M-F444W$ color indicates strong CO$_2$ absorption in its atmosphere and suggests significant enhancement in metals compared to free-floating objects. From the ATMO hot-start evolutionary models, we estimate an effective temperature of $600^{+45}_{-60}$ K and mass of $2.4\pm0.6$ $M_{\rm Jup}$, which also closely matches similar estimates for 51 Eri b. $β$ Pic d is among the lowest-mass exoplanets imaged from the ground. This discovery highlights the deep sensitivity achievable with ground-based imaging in the mid-infrared and the discovery potential of future high-contrast observations with the Extremely Large Telescope.

astro-ph.EP

Polynomial encoding of rooted trees with branch lengths

Phylogenetic trees are rooted trees with branch lengths that record genetic divergence or elapsed time, and quantifying differences between them is central to a wide range of evolutionary and epidemiological analyses. Graph-polynomial encodings of rooted trees provide an accurate, interpretable, and computationally efficient way to compare tree shapes, but existing polynomial encodings must be paired with auxiliary structures to study rooted trees with branch lengths. We introduce a bivariate polynomial encoding that incorporates branch lengths directly into a recursive computation from the leaf vertices to the root vertex of a tree. We prove that, for rooted trees with branch lengths and no vertices of degree two, which include all standard phylogenetic trees, two trees have the same polynomial if and only if their underlying unlabeled trees are isomorphic and the branch lengths of corresponding edges are equal. We apply the polynomial encoding to three published HIV-1 phylogenies sampled in different epidemiological settings and show that it accurately separates the three datasets based on their tree topologies and branch lengths, outperforming previous polynomial-based approaches for analyzing rooted trees with branch lengths.

q-bio.PE

A New Multi-Domain Benchmark for Micro-Action Recognition and Detection

Micro-actions are short-duration, low-amplitude subtle body movements at the whole-body level that can reveal latent intentions, involuntary reactions, and fine-grained affective changes. Our previous MA-52 benchmark has provided an important foundation for micro-action recognition, but it remains limited in scale, scene diversity, task coverage, and evaluation protocols. To advance micro-action analysis toward more realistic and comprehensive settings, we introduce MMA-82, a large-scale multi-domain extension of MA-52. MMA-82 expands the label space from 52 to 82 fine-grained micro-action categories and covers four distinct domains, including laboratory interviews, street interviews, psychiatric patient interviews, and emotion-rich television videos, resulting in 77,856 annotated instances from 454 subjects. Built upon MMA-82, we establish two core tasks: Micro-Action Recognition and Multi-label Micro-Action Detection. For recognition, we further define in-domain and cross-domain protocols, including few-shot and zero-shot settings, to evaluate model robustness, transferability, and generalization. Extensive experiments show that current methods still struggle with realistic micro-action understanding, especially under domain shift, long-tailed category distributions, and complex temporal localization. Beyond benchmarking, we investigate the relationship between micro-actions and emotion, showing that micro-actions are strongly associated with emotional states and provide complementary cues to facial micro-expressions for improved emotion recognition. These results demonstrate that MMA-82 serves as a comprehensive and challenging benchmark for realistic micro-action analysis and a valuable resource for human-centered AI. MMA-82 is available at https://lpynow.github.io/MMA-82-AIM/.

cs.CV

QALM: Escaping Local Minima via Interleaved Exploration and Exploitation in Quantum Circuit Optimization

Quantum circuit optimizers face a fundamental limitation in how they tolerate temporary cost increases. At one extreme, greedy rule-based optimizers immediately apply any cost-reducing transformation, achieving high efficiency but quickly becoming trapped in local minima. At the other extreme, search-based optimizers accept cost-increasing moves to explore the circuit space and escape such minima. However, because search-based optimizers cannot determine within a reasonable time budget whether a given point is promising, that is, whether its neighborhood contains a deeper local minimum, they must blindly explore higher-cost regions. As a result, escaping the current basin to reach a promising point takes exponentially many steps. In this work, we show that this limitation can be overcome with a hybrid framework that interleaves the exhaustive exploration capabilities of search algorithms with the efficiency of rule-based optimization. We implement this framework as QALM, a novel optimizer designed to escape local minima without incurring the runtime penalties of pure search. Crucially, our results demonstrate that QALM does not merely strike a balance; it outperforms existing rule-based and search-based optimizers in circuit reduction rates while operating with the computational efficiency of rule-based systems. In a comprehensive evaluation across 248 circuits, QALM matches or exceeds the fidelity of the strongest baseline on 83.9% of these circuits, given the same time budget.

quant-ph

A Multi-Modal Framework with Cross-Subject Pseudo-Labeling and Semantic Alignment for Micro-Gesture Recognition

Micro-gestures (MGs) are spontaneous and subtle body movements that frequently convey hidden human emotions. Recognizing MGs in untrimmed videos remains highly challenging due to their extremely low signal-to-noise ratio, severe long-tailed class distribution, and the inherent domain shift encountered in cross-subject evaluation scenarios. In this paper, we propose a comprehensive multi-modal framework for Track 1 of the 4th MiGA-IJCAI Challenge. To capture fine-grained representations, we design a saliency-guided multi-modal extraction pipeline integrating 68-keypoint skeleton joint coordinates, 3D heatmap volumes, and high-resolution RGB visual features. We introduce a gentle square-root smoothed weighting mechanism paired with an Orthogonal Semantic Embedding Loss to protect tail classes without compromising overall recognition capabilities. More importantly, to bridge the cross-subject generalization gap, we propose a Cross-Modal Pseudo-Labeling (CMPL) strategy for unsupervised domain adaptation, which significantly boosts single-modal robustness. A temperature-scaled soft-voting mechanism is finally utilized to alleviate overconfidence during late fusion. Extensive experiments demonstrate that our framework achieves a competitive F1-score of 68.13\%, securing the 4th place.

cs.CV

Measuring language complexity from hierarchical reuse of recurring patterns

We introduce the ladderpath index as a measure of language complexity grounded in algorithmic information theory. It counts the minimum steps needed to reconstruct a sequence through hierarchical reuse of repeated substructures, capturing an exactly computable but constrained form of algorithmic compressibility related to, but distinct from, Kolmogorov complexity. We apply the ladderpath approach to 21 parallel corpora from the Parallel Universal Dependencies dataset. The ladderpath index is approximately invariant across the languages, and varies much less than the corpus length. This is more pronounced when all corpora are mapped to a unified binary representation, providing evidence for the equi-complexity hypothesis from a representation-independent perspective. We also observe trade-offs between character inventory size and corpus length, and between vocabulary-level and corpus-level reconstruction complexity, supporting the trade-off hypothesis that total complexity is conserved and redistributed across linguistic levels. The reusable substructures identified by the ladderpath approach, without any linguistic input, overlap with words and morphological components attested in the natural vocabulary. The hierarchical reuse captured by the ladderpath approach parallels the chunking mechanisms proposed in cognitive science, where the human cognitive system compresses linguistic input into nested, reusable units under shared memory and processing constraints. This connection between cognitive chunking and the ladderpath approach provides a new interpretation for the equi-complexity and trade-off hypotheses, grounding both in the shared cognitive architecture that underlies language processing across human languages.

cs.CL

Motion Reinforces Appearance: RGB-Skeleton Gated Residual Fusion for Micro-Gesture Online Recognition

Micro-gesture analysis attracts increasing attention for inferring spontaneous emotion from subtle body movements. Micro-gesture online recognition, which localizes and classifies each gesture instance in untrimmed videos, is a core task in the 4th EI-MiGA-IJCAI Challenge. Compared with typical temporal action detection, MGR emphasizes the localization and classification of actions, requiring the model to output the start time, end time, and category of each micro-gesture. Moreover, since micro-gestures are highly spontaneous, relying solely on a single modality makes it difficult to capture the complete and accurate multi-modal cues. In this work, we propose DyFADet+, which extends DyFADet into a dual-stream RGB-skeleton framework. In our model, both modalities are projected into shared multi-scale temporal embeddings and fused through a gated residual module, which adaptively injects skeleton motion into the RGB representation rather than using naive concatenation. Finally, these fused features are decoded by a Dynamic TAD head for online classification and boundary regression. On the SMG dataset, our method achieves an F1 score of 40.88, ranking 2nd in the Micro-gesture Online Recognition track.

cs.CV

Achieving Optimal-Distance Atom-Loss Correction via Pauli Envelope

Atom loss is a major error source in neutral-atom quantum computers, accounting for over 40% of the total physical errors in recent experiments. Its nonlinear and correlated nature poses significant challenges: current syndrome extraction circuits require additional overhead or sacrifice loss tolerance, and existing decoders are computationally inefficient, suboptimal, or lack provable guarantees. To address these challenges, we propose the Pauli Envelope framework, which bounds the effect of atom loss with low-weight, efficiently computable Pauli approximations, generalizing existing loss-to-Pauli methods and enabling rigorous analysis. Guided by this framework, we design improved atom-replenishing syndrome extraction circuits, the Mid-SWAP syndrome extraction, which achieves optimal loss distance and minimal space-time overhead for rotated surface codes. We also propose two decoders: an Envelope-MLE decoder achieving the optimal loss distance d_loss ~ d, and an Envelope-Matching decoder achieving d_loss ~ 2d/3 via Minimum-Weight Perfect Matching (MWPM), surpassing the previous best (d_loss ~ d/2) and readily integrating with fast correlated decoding techniques for transversal logical circuits. Circuit-level simulations demonstrate up to 40% higher thresholds and 30% higher effective distances compared with existing methods in the loss-dominated regime. Moreover, we explore correlated atom loss and show that it is easier to correct than independent loss, with thresholds rising from 5.15% to 7.82%. Remarkably, our Envelope-MLE decoder improves the error suppression factor of a hybrid MLE--machine-learning decoder from Λ= 2.14 to Λ= 2.24 on recent experimental data.

quant-ph

Generalized matching decoders for 2D topological translationally-invariant codes

Two-dimensional topological translationally-invariant (TTI) quantum codes, such as the toric code (TC) and bivariate bicycle (BB) codes, are promising candidates for fault-tolerant quantum computation. For such codes to be practically relevant, their decoders must successfully correct the most likely errors while remaining computationally efficient. For the TC, graph-matching decoders satisfy both requirements and, additionally, admit provable performance guarantees. Given the equivalence between TTI codes and (multiple copies of) the TC, one may then ask whether TTI codes also admit analogous graph-matching decoders. In this work, we develop a graph-matching approach to decoding general TTI codes. Intuitively, our approach coarse-grains the TTI code to obtain an effective description of the syndrome in terms of TC excitations, which can then be removed using graph-matching techniques. We prove that our decoders correct errors of weight up to a constant fraction of the code distance and achieve non-zero code-capacity thresholds. We further numerically study a variant optimized for practically relevant BB codes and observe performance comparable to that of the belief propagation with ordered statistics decoder. Our results indicate that graph-matching decoders are a viable approach to decoding BB codes and other TTI codes.

quant-ph

DualSentinel: A Lightweight Framework for Detecting Targeted Attacks in Black-box LLM via Dual Entropy Lull Pattern

Recent intelligent systems integrate powerful Large Language Models (LLMs) through APIs, but their trustworthiness may be critically undermined by targeted attacks like backdoor and prompt injection attacks, which secretly force LLMs to generate specific malicious sequences. Existing defensive approaches for such threats typically rely on high access rights, impose prohibitive costs, and hinder normal inference, rendering them impractical for real-world scenarios. To solve these limitations, we introduce DualSentinel, a lightweight and unified defense framework that can accurately and promptly detect the activation of targeted attacks alongside the LLM generation process. We first identify a characteristic of compromised LLMs, termed Entropy Lull: when a targeted attack successfully hijacks the generation process, the LLM exhibits a distinct period of abnormally low and stable token probability entropy, indicating it is following a fixed path rather than making creative choices. DualSentinel leverages this pattern by developing an innovative dual-check approach. It first employs a magnitude and trend-aware monitoring method to proactively and sensitively flag an entropy lull pattern at runtime. Upon such flagging, it triggers a lightweight yet powerful secondary verification based on task-flipping. An attack is confirmed only if the entropy lull pattern persists across both the original and the flipped task, proving that the LLM's output is coercively controlled. Extensive evaluations show that DualSentinel is both highly effective (superior detection accuracy with near-zero false positives) and remarkably efficient (negligible additional cost), offering a truly practical path toward securing deployed LLMs. The source code can be accessed at https://doi.org/10.5281/zenodo.18479273.

cs.CR

A Comparison of Polynomial-Based Tree Clustering Methods

Tree structures appear in many fields of the life sciences, including phylogenetics, developmental biology and nucleic acid structures. Trees can be used to represent RNA secondary structures, which directly relate to the function of non-coding RNAs. Recent developments in sequencing technology and artificial intelligence have yielded numerous biological data that can be represented with tree structures. This requires novel methods for tree structure data analytics. Tree polynomials provide a computationally efficient, interpretable and comprehensive way to encode tree structures as matrices, which are compatible with most data analytics tools. Machine learning methods based on the Canberra distance between tree polynomials have been introduced to analyze phylogenies and nucleic acid structures. In this paper, we compare the performance of different distances in tree clustering methods based on a tree distinguishing polynomial. We also implement two basic autoencoder models for clustering trees using the polynomial. We find that the distance based methods with entry-level normalized distances have the highest clustering accuracy among the compared methods.

cs.LG

Sensitivity to Sub-Io-sized Exosatellite Transits in the MIRI LRS Lightcurve of the Nearest Substellar Worlds

JWST's unprecedented sensitivity enables precise spectrophotometric monitoring of substellar worlds, revealing atmospheric variability driven by mechanisms operating across different pressure levels. This same precision now permits exceptionally sensitive searches for transiting exosatellites, small terrestrial companions to these worlds. Using a novel simultaneous dual-band search method to address host variability, we present a search for transiting exosatellites in an 8-hour JWST/MIRI LRS lightcurve of the nearby ($2.0\,pc$) substellar binary WISE J1049-5319AB, composed of two $\sim30 M_{\rm Jup}$ brown dwarfs separated by $3.5\,au$ and viewed near edge-on. Although we detect no statistically significant transits, our injection-recovery tests demonstrate sensitivity to satellites as small as $0.275\,R_{\oplus}$ ($0.96\,R_{\rm Io}$ or $\sim$1 lunar radius), corresponding to 300ppm transit depths, and satellite-to-host mass ratios $>$$10^{-6}$. This approach paves the way for detecting Galilean-moon analogs around directly imaged brown dwarfs, free-floating planets, and wide-orbit exoplanets, dozens of which are already scheduled for JWST lightcurve monitoring. In our Solar System, each giant planet hosts on average 3.5 moons above this threshold, suggesting that JWST now probes a regime where such companions are expected to be abundant. The technique and sensitivities demonstrated here mark a critical step toward detecting exosatellites and ultimately enabling constraints on the occurrence rates of small terrestrial worlds orbiting $1\text{-}70$$M_{\rm Jup}$ hosts.

astro-ph.EP

Online Micro-gesture Recognition Using Data Augmentation and Spatial-Temporal Attention

In this paper, we introduce the latest solution developed by our team, HFUT-VUT, for the Micro-gesture Online Recognition track of the IJCAI 2025 MiGA Challenge. The Micro-gesture Online Recognition task is a highly challenging problem that aims to locate the temporal positions and recognize the categories of multiple micro-gesture instances in untrimmed videos. Compared to traditional temporal action detection, this task places greater emphasis on distinguishing between micro-gesture categories and precisely identifying the start and end times of each instance. Moreover, micro-gestures are typically spontaneous human actions, with greater differences than those found in other human actions. To address these challenges, we propose hand-crafted data augmentation and spatial-temporal attention to enhance the model's ability to classify and localize micro-gestures more accurately. Our solution achieved an F1 score of 38.03, outperforming the previous state-of-the-art by 37.9%. As a result, our method ranked first in the Micro-gesture Online Recognition track.

cs.CV

MMAD: Multi-label Micro-Action Detection in Videos

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising applications in human emotion analysis. In real-world scenarios, human micro-actions often temporally co-occur, with multiple micro-actions overlapping in time, such as concurrent head and hand movements. However, current research primarily focuses on recognizing individual micro-actions while overlooking their co-occurring nature. To address this gap, we propose a new task named Multi-label Micro-Action Detection (MMAD), which involves identifying all micro-actions in a given short video, determining their start and end times, and categorizing them. Accomplishing this requires a model capable of accurately capturing both long-term and short-term action relationships to detect multiple overlapping micro-actions. To facilitate the MMAD task, we introduce a new dataset named Multi-label Micro-Action-52 (MMA-52) and propose a baseline method equipped with a dual-path spatial-temporal adapter to address the challenges of subtle visual change in MMAD. We hope that MMA-52 can stimulate research on micro-action analysis in videos and prompt the development of spatio-temporal modeling in human-centric video understanding. The proposed MMA-52 dataset is available at: https://github.com/VUT-HFUT/Micro-Action.

cs.CV

ConiQ: Enabling Concatenated Quantum Error Correction on Neutral Atom Arrays

Recent progress on concatenated codes, especially many-hypercube codes, achieves unprecedented space efficiency. Yet two critical challenges persist in practice. First, these codes lack efficient implementations of addressable logical gates. Second, the required high degree of parallelism and long-range interactions pose significant challenges for current hardware platforms. In this paper, we propose an efficient compilation approach for concatenated codes, specifically many-hypercube codes, targeted at neutral atom arrays, which provide the necessary parallelism and long-range interactions. Our approach builds on two key innovations. First, we introduce Automorphism-assisted Hierarchical Addressing (AHA) logical CNOT gates that significantly reduce spacetime overhead compared to conventional distillation-based methods. Second, we develop Virtual Atom Intermediate Representation (VAIR) that enables level-wise optimization and legalization. We implement these innovations in ConiQ, a hardware-aware quantum compiler designed to compile fault-tolerant quantum circuits for neutral atom arrays using many-hypercube codes. Our evaluation demonstrates that ConiQ achieves up to 2000x reduction in spacetime overhead and up to 10^6x reduction in compilation time compared to state-of-the-art compilers, with our AHA gates providing an additional overhead reduction of up to 20x. These results establish concatenated codes as a promising approach for fault-tolerant quantum computing in the near future.

cs.AR