SearcharxivSearch

arXiv subjects

Yiqi Liu

Publications and source records attributed to Yiqi Liu.

At least 19 recordsLinked to original sources

Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with VOXEL

To overcome the well-known memory bottleneck of AI chips, 3D-stacked architectures that employ advanced packaging technology with high-density through-silicon vias (TSVs) pins have proven to be a promising solution. The 3D-stacked AI chip enables ultra-high memory bandwidth between compute and memory by stacking numerous DRAM banks atop many AI cores in a distributed manner. However, it is not easy to explore the efficiency of the 3D-stacked AI chip, due to its unique distributed nature. And we need to carefully consider multiple intertwined factors that range from upper-level computing paradigm to machine learning (ML) compiler optimizations, and to the underlying hardware architecture. In this paper, we develop VOXEL, a fast and compiler-aware end-to-end simulation framework to facilitate exploring the efficiency of 3D-stacked AI chips for large language model (LLM) inference. VOXEL enables the software/hardware co-exploration by employing a programming interface that allows ML compilers to customize the model execution plans. After validating the results of VOXEL with an emulator on real silicon, we thoroughly examine the impact and correlation of different aspects of 3D-stacked AI chips, including state-of-the-art compute paradigms, tile-to-core mapping, tensor-to-bank mapping, NoC topologies and link bandwidth, DRAM bank bandwidth, per-core SRAM capacity, and energy/thermal constraints. Our findings disclose that the end-to-end efficiency of a 3D stacked AI chip not only is determined by the cooperative function of these factors, but also significantly depends on the mappings from tiles to AI core and DRAM banks. We report our findings in the paper, expecting that they will shed light on the development of 3D-stacked AI chip ecosystem. We open source VOXEL for public research.

cs.AR

When Tool-Backed Skill Retrieval Fails: Source-Style Collapse in Executable Capability Retrieval

Large-scale agents increasingly rely on retrieval to access external capabilities. We study this retrieval gate in structured tools and APIs, a measurable class of tool-backed executable skills that must be surfaced before an agent can plan, incorporate, or act. In this setting the retrieval layer can silently fail even when the capability corpus is fixed: on ToolRet, a retriever fine-tuned on one source-specific slice collapses on another source-specific slice of the same benchmark, with FT-1100 despite its higher lexical overlap with the gold tools. We call this failure mode source-style collapse. Query-side TF-IDF fingerprints flag source styles on which the fine-tuned retriever is likely to fail better than semantic or length-based proxies, giving a cheap signal for mismatch over a fixed tool corpus. We propose ToolScout, a source-aware routing method that uses this signal as a routing guard: on the mixed 4,996-query stream, TF-IDF-based routing raises coverage from 22.3% to 86.1%, and across five collapsed sources 20 matched examples raise the coverage-weighted global top-1 proxy from 1.3% to 53.9%. The same failure and routing behaviors persist when tools are rerendered as executable skill cards, which rules out raw API-schema format as the sole cause.

cs.LG

Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration

When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or merely suppressed at the output layer. We investigate this mechanism using a controlled withhold setting, which yields perfectly matched answering and refusal trajectories for bidirectional activation patching. We uncover a causal asymmetry in intervention locality under matched causal interventions, which we term broken symmetry. Even when a model generates a clean refusal, the correct answer remains linearly recoverable from its hidden states. Furthermore, releasing this withheld answer is a highly local operation, requiring only a single-position patch. Conversely, the reverse operation is not equally local: reimposing suppression requires broader interventions across multiple positions, and assembling a coherent refusal sequence is more difficult still. We further demonstrate that while an average answer-to-refusal displacement vector marks the geometric difference between these states, it fails to act as a reliable, reversible linear control toggle between behaviours. Taken together, our findings show that refusal does not function as a simple symmetric switch. For safety and auditing, this implies that probe recoverability can overestimate true behavioural control, and locating refusal-relevant directions does not reliably grant the ability to steer a model from answering to coherent refusal.

cs.AI

Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos

Social media videos often communicate meanings that go beyond their visible actions, captions, or speech. A mundane clip may become humorous, ironic, or satire only through the interaction of multimodal cues and cultural context, making such content a difficult test case for video-language models. In this paper, we introduce \textit{DrivelHub+}, a benchmark for evaluating whether models can infer the implicit, non-linear, and rhetorically layered meanings of social media videos that appear nonsensical on the surface but convey deliberate pragmatic meanings. DrivelHub+ consists of 1,000 videos collected from social media, each annotated with a human-written implicit narrative explanation. Unlike conventional video understanding tasks focused on recognition or description, we present a benchmark that targets contextual multimodal reasoning. We evaluate current video-language models from two perspectives: explanation, where models must explain the pragmatic comprehension of a video in natural language; and representation, where we adapt reasoning-as-retrieval to test whether model representations align videos with their corresponding implicit narratives in both video-to-text and text-to-video retrieval. Our benchmark provides a diagnostic setting for measuring the gap between multimodal perception and pragmatic comprehension, asking whether current models can move beyond describing what is shown to inferring what is meant.

cs.CL

PerturbMap: Cross-Context Transfer of Single-Cell Perturbation Responses

Single-cell perturbation atlases rarely measure every intervention in every cellular context: a query perturbation is often observed in one or more source contexts but missing in the recipient context where its effect is needed. Ignoring those measured responses discards query-specific experimental evidence, whereas copying or weakly calibrating them across contexts risks transferring the wrong signal. We propose PerturbMap, which predicts a missing recipient-context effect by combining a recipient-local low-rank base with accepted proposals that transport the same perturbation's measured source responses through source-to-recipient ridge experts fit on paired training perturbations, with proposal weights determined by route reliability estimated on validation anchors. On the Perturb-CITE-seq melanoma cohort, PerturbMap improves full-effect MSE by 4.1\% over a recipient-local low-rank base and achieves lower MSE than FedAvg, zero-response, raw-copy, calibrated-copy, and identity-shuffled affine controls. It remains within $2.82\times10^{-6}$ MSE of our centralized token-matched pooled reference, which uses a stronger training interface. A condition-mean specificity diagnostic shows the same direction: same-recipient top-10 counterpart retrieval by cosine increases from 74.5\% for the low-rank base to 80.5\% for PerturbMap.

cs.AI

Identification and Inference for Algorithmic Frontiers with Selective Labels

This paper provides identification results to characterize a fairness-accuracy (FA) frontier, and statistical inference tools to test hypotheses and build a confidence set for the FA-frontier, when outcomes are observed only for selected individuals. When the selection process is unrestricted but loss is measured in specific ways, we provide a characterization of the sharp identification region of the FA-frontier. Under an assumption of unconfoundedness conditional on observables (and unrestricted loss functions), we obtain point identification and propose a debiased machine learning estimator, derive its asymptotic distribution, and show how this can be used to carry out inference for the FA-frontier. In work in progress, we extend the partial identification results to a broader class of loss functions.

econ.EM

Calibration of CMB Polarisation Using Cross-Experiment Correlations

Parity-violating physics in the Universe can generate correlations between the Cosmic Microwave Background (CMB) $E$- and $B$-modes, but detecting such signals requires extremely accurate calibration of instruments. We describe a data-driven method to calibrate the relative polarisation angle between CMB experiments using cross-correlations of observations over a common sky region. Unlike standard self-calibration approaches, this method does not assume vanishing isotropic cosmic birefringence or primordial $EB$ correlations when estimating the relative misalignment angle, and therefore preserves sensitivity to parity-violating physics. As a proof of concept, we forecast the performance of this method using the Simons Observatory (SO) Small Aperture Telescopes (SATs) as a calibrated reference. If they can be calibrated to an uncertainty of $0.08^\circ$, as anticipated from the SO wire grid calibration system, we show that the SO Large Aperture Telescope and Planck could be calibrated to uncertainties of $0.10^\circ$ and $0.17^\circ$, respectively, at $\sim 145$ GHz. This approach relies on the availability of at least one well-calibrated instrument, and provides a complementary path to improving polarisation calibration across experiments, enabling more robust searches for parity-violating physics in the CMB, such as cosmic birefringence.

astro-ph.CO

Closed-Form Analytical Charge Response Model for Silicon Photomultipliers with Recursive Correlated Avalanches

Silicon photomultipliers (SiPMs) have become the preferred photodetectors in next-generation neutrino experiments, yet no unified closed-form analytical expression free of truncation and numerical convolution has been established for their full charge response spectrum, which must simultaneously capture correlated cross-talk and afterpulsing effects absent in conventional photomultiplier tubes (PMTs). We present a unified closed-form model for the SiPM charge response within the characteristic-function framework, treating pedestal noise, single-electron-response (SER) charge, internal optical cross-talk, and afterpulsing on equal footing. The characteristic-function representation factorises the full charge spectrum into three independent physical components: pedestal, single-electron response (SER), and avalanche count statistics. Prompt internal optical cross-talk is modelled as a Galton-Watson branching process with Poisson offspring; building on the Generalised Poisson count statistics identified by Vinogradov, we derive a Lambert $W$ closed form for the total-progeny PGF via Lagrange-Bürmann inversion, providing the analytical handle needed for efficient event-level reconstruction. Afterpulsing is modelled as a per-avalanche geometric chain, derived as the maximum-entropy Poisson-Gamma mixture: the exponential prior-maximum-entropy for a positive continuous yield with fixed mean-marginalised over a Poisson count yields the geometric per-avalanche distribution, whose $N$-avalanche total is Negative Binomial. This naturally encompasses the Poisson afterpulsing limit and recursive afterpulse chains while preserving analytical closure. The resulting eight-parameter expression is further applied to derive an explicit per-channel charge-time likelihood for event-level energy reconstruction without numerical convolution at inference time.

physics.ins-det

Ouroboros: Wafer-Scale SRAM CIM with Token-Grained Pipelining for Large Language Model Inference

Conventional LLM inference architectures suffer from high energy and latency due to frequent data movement across memory hierarchies. We propose Ouroboros, a wafer-scale SRAM-based Computing-in-Memory (CIM) architecture that executes all operations in situ, eliminating off-chip migration. To maximize its limited first-level capacity, we introduce three innovations: Token-Grained Pipelining: Replaces sequence-level pipelining to mitigate length variations, boosting utilization and reducing activation storage. Distributed Dynamic KV Cache Management: Decouples memory from compute to leverage fragmented SRAM for efficient KV storage. Communication-Aware Mapping: Optimizes core allocation for locality and fault tolerance across the wafer. Experimental results show Ouroboros achieves average gains of $4.1\times$ in throughput and $4.2\times$ in energy efficiency, peaking at $9.1\times$ and $17\times$ for the 13B model. (*Due to the notification of arXiv "The Abstract field cannot be longer than 1,920 characters", the appeared Abstract is shortened. For the full Abstract, please download the Article.)

cs.AR

The Achilles' Heel of Angular Margins: A Chebyshev Polynomial Fix for Speaker Verification

Angular margin losses, such as AAM-Softmax, have become the de facto in speaker and face verification. Their success hinges on directly manipulating the angle between features and class prototypes. However, this manipulation relies on the arccos function to recover the angle, introducing a significant yet overlooked source of training instability. The derivative of arccos explodes at its boundaries, causing gradient peaks during optimisation. Furthermore, the formulation fails to generate a sufficiently sharp gradient for hard-to-classify examples. We address these issues by proposing ChebyAAM, a loss that replaces the arccos operation with its Chebyshev polynomial approximation. This substitution eliminates gradient explosion and applies a stronger corrective signal to hard examples, leading to more effective optimisation. Experiments on three benchmarks (VoxCeleb, SITW, and CN-Celeb) demonstrate that our method resolves the instability and consistently improves performance. Our work suggests that approximating angular operations, rather than calculating them explicitly, offers a more robust path for designing future metric learning losses. Code is available at https://github.com/ExtraOrdinaryLab/vibe.

cs.SD

ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation

Evaluating the quality of generated text automatically remains a significant challenge. Conventional reference-based metrics have been shown to exhibit relatively weak correlation with human evaluations. Recent research advocates the use of large language models (LLMs) as source-based metrics for natural language generation (NLG) assessment. While promising, LLM-based metrics, particularly those using smaller models, still fall short in aligning with human judgments. In this work, we introduce ContrastScore, a contrastive evaluation metric designed to enable higher-quality, less biased, and more efficient assessment of generated text. We evaluate ContrastScore on two NLG tasks: machine translation and summarization. Experimental results show that ContrastScore consistently achieves stronger correlation with human judgments than both single-model and ensemble-based baselines. Notably, ContrastScore based on Qwen 3B and 0.5B even outperforms Qwen 7B, despite having only half as many parameters, demonstrating its efficiency. Furthermore, it effectively mitigates common evaluation biases such as length and likelihood preferences, resulting in more robust automatic evaluation.

cs.CL

Synthetic Parallel Trends

Popular empirical strategies for policy evaluation in the panel data literature -- including difference-in-differences (DID), synthetic control (SC) methods, and their variants -- rely on key identifying assumptions that can be expressed through a specific choice of weights $ω$ relating pre-treatment trends to the counterfactual outcome. While each choice of $ω$ may be defensible in empirical contexts that motivate a particular method, it relies on fundamentally untestable and often fragile assumptions. I develop an identification framework that allows for all weights satisfying a Synthetic Parallel Trends assumption: the treated unit's trend is parallel to a weighted combination of control units' trends for a general class of weights. The framework nests these existing methods as special cases and is by construction robust to violations of their respective assumptions. I construct a valid confidence set for the identified set of the treatment effect, which admits a linear programming representation with estimated coefficients and nuisance parameters that are profiled out. In simulations where the assumptions underlying DID or SC-based methods are violated, the proposed confidence set remains robust and attains nominal coverage, while existing methods suffer severe undercoverage.

econ.EM

Toward Engineering AGI: Benchmarking the Engineering Design Capabilities of LLMs

Modern engineering, spanning electrical, mechanical, aerospace, civil, and computer disciplines, stands as a cornerstone of human civilization and the foundation of our society. However, engineering design poses a fundamentally different challenge for large language models (LLMs) compared with traditional textbook-style problem solving or factual question answering. Although existing benchmarks have driven progress in areas such as language understanding, code synthesis, and scientific problem solving, real-world engineering design demands the synthesis of domain knowledge, navigation of complex trade-offs, and management of the tedious processes that consume much of practicing engineers' time. Despite these shared challenges across engineering disciplines, no benchmark currently captures the unique demands of engineering design work. In this work, we introduce EngDesign, an Engineering Design benchmark that evaluates LLMs' abilities to perform practical design tasks across nine engineering domains. Unlike existing benchmarks that focus on factual recall or question answering, EngDesign uniquely emphasizes LLMs' ability to synthesize domain knowledge, reason under constraints, and generate functional, objective-oriented engineering designs. Each task in EngDesign represents a real-world engineering design problem, accompanied by a detailed task description specifying design goals, constraints, and performance requirements. EngDesign pioneers a simulation-based evaluation paradigm that moves beyond textbook knowledge to assess genuine engineering design capabilities and shifts evaluation from static answer checking to dynamic, simulation-driven functional verification, marking a crucial step toward realizing the vision of engineering Artificial General Intelligence (AGI).

cs.CE

Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts

Evaluating natural language generation systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, lack of standardisation, and demographic biases, limiting reproducibility. LLM-based evaluators offer a scalable alternative but are highly sensitive to prompt design, where small variations can lead to significant discrepancies. In this work, we propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions, enabling the automatic generation of highly effective, model-specific evaluation prompts. Our method requires only a single evaluation sample and eliminates the need for time-consuming manual prompt engineering, thereby improving both efficiency and robustness. Our work contributes toward a new direction for more robust and efficient LLM-based evaluation.

cs.CL

ELK: Exploring the Efficiency of Inter-core Connected AI Chips with Deep Learning Compiler Techniques

To meet the increasing demand of deep learning (DL) models, AI chips are employing both off-chip memory (e.g., HBM) and high-bandwidth low-latency interconnect for direct inter-core data exchange. However, it is not easy to explore the efficiency of these inter-core connected AI (ICCA) chips, due to a fundamental tussle among compute (per-core execution), communication (inter-core data exchange), and I/O (off-chip data access). In this paper, we develop Elk, a DL compiler framework to maximize the efficiency of ICCA chips by jointly trading off all the three performance factors discussed above. Elk structures these performance factors into configurable parameters and forms a global trade-off space in the DL compiler. To systematically explore this space and maximize overall efficiency, Elk employs a new inductive operator scheduling policy and a cost-aware on-chip memory allocation algorithm. It generates globally optimized execution plans that best overlap off-chip data loading and on-chip execution. To examine the efficiency of Elk, we build a full-fledged emulator based on a real ICCA chip IPU-POD4, and an ICCA chip simulator for sensitivity analysis with different interconnect network topologies. Elk achieves 94% of the ideal roofline performance of ICCA chips on average, showing the benefits of supporting large DL models on ICCA chips. We also show Elk's capability of enabling architecture design space exploration for new ICCA chip development.

cs.AR

The Simons Observatory: Assessing the Impact of Dust Complexity on the Recovery of Primordial $B$-modes

We investigate how dust foreground complexity can affect measurements of the tensor-to-scalar ratio, $r$, in the context of the Simons Observatory, using a cross-spectrum component separation analysis. Employing a suite of simulations with realistic Galactic dust emission, we find that spatial variation in the dust frequency spectrum, parametrized by $β_d$, can bias the estimate for $r$ when modeled using a low-order moment expansion to capture this spatial variation. While this approach performs well across a broad range of dust complexity, the bias increases with more extreme spatial variation in dust frequency spectrum, reaching as high as $r\sim0.03$ for simulations with no primordial tensors and a spatial dispersion of $σ(β_d)\simeq0.3$ -- the most extreme case considered, yet still consistent with current observational constraints. This bias is driven by changes in the $\ell$-dependence of the dust power spectrum as a function of frequency that can mimic a primordial $B$-mode tensor signal. Although low-order moment expansions fail to capture the full effect when the spatial variations of $β_d$ become large and highly non-Gaussian, our results show that extended parametric methods can still recover unbiased estimates of $r$ under a wide range of dust complexities. We further find that the bias in $r$, at the highest degrees of dust complexity, is largely insensitive to the spatial structure of the dust amplitude and is instead dominated by spatial correlations between $β_d$ and dust amplitude, particularly at higher orders. If $β_d$ does spatially vary at the highest levels investigated here, we would expect to use more flexible foreground models to achieve an unbiased constraint on $r$ for the noise levels anticipated from the Simons Observatory.

astro-ph.CO

Inference for an Algorithmic Fairness-Accuracy Frontier

Algorithms are increasingly used to aid with high-stakes decision making. Yet, their predictive ability frequently exhibits systematic variation across population subgroups. To assess the trade-off between fairness and accuracy using finite data, we propose a debiased machine learning estimator for the fairness-accuracy frontier introduced by Liang, Lu, Mu, and Okumura (2024). We derive its asymptotic distribution and propose inference methods to test key hypotheses in the fairness literature, such as (i) whether excluding group identity from use in training the algorithm is optimal and (ii) whether there are less discriminatory alternatives to a given algorithm. In addition, we construct an estimator for the distance between a given algorithm and the fairest point on the frontier, and characterize its asymptotic distribution. Using Monte Carlo simulations, we evaluate the finite-sample performance of our inference methods. We apply our framework to re-evaluate algorithms used in hospital care management and show that our approach yields alternative algorithms that lie on the fairness-accuracy frontier, offering improvements along both dimensions.

econ.EM

Using Forests in Multivariate Regression Discontinuity Designs

We discuss estimation and inference of conditional treatment effects in regression discontinuity (RD) designs with multiple scores. In addition to local linear regressions and the minimax-optimal estimator more recently proposed by Imbens and Wager (2019), we argue that two variants of random forests, honest regression forests and local linear forests, should be added to the toolkit of applied researchers working with multivariate RD designs; their validity follows from results in Wager and Athey (2018) and Friedberg et al. (2020). We design a systematic Monte Carlo study with data generating processes built both from functional forms that we specify and from Wasserstein Generative Adversarial Networks that closely mimic the observed data. We find no single estimator dominates across all specifications: (i) local linear regressions perform well in univariate settings, but the common practice of reducing multivariate scores to a univariate one can incur under-coverage, possibly due to vanishing density at the transformed cutoff; (ii) good performance of the minimax-optimal estimator depends on accurate estimation of a nuisance parameter and its current implementation only accepts up to two scores; (iii) forest-based estimators are not designed for estimation at boundary points and are susceptible to finite-sample bias, but their flexibility in modeling multivariate scores opens the door to a wide range of empirical applications, as illustrated by an empirical study of COVID-19 hospital funding with three eligibility criteria.

econ.EM