SearcharxivSearch

arXiv subjects

Yifan Cao

Publications and source records attributed to Yifan Cao.

At least 19 recordsLinked to original sources

Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT

Translating C code into safe, idiomatic Rust is a longstanding software-engineering goal because it can eliminate entire classes of memory-safety vulnerabilities while preserving the functional behavior of legacy systems. Large language models (LLMs) have shown promise for this task but typically underperform when applied off-the-shelf, since general-purpose pretraining rarely emphasizes idiomatic Rust generation, cross-language semantic equivalence, or the ability to reason about and repair compiler/runtime feedback. In this report we describe a three-stage fine-tuning curriculum applied to Qwen3-27B that is designed to progressively specialize the model for the C-to-Rust (C2Rust) translation task: (1) continued pretraining on Rust-centric corpora to strengthen the model's prior over idiomatic Rust syntax and standard-library usage; (2) supervised fine-tuning (SFT) on the microsoft/Verus_Training_Data dataset to instill debugging and self-repair behavior over Rust code; and (3) task-specific SFT on paired C/Rust solutions derived from LeetCode problems to teach direct semantic translation. We evaluate the resulting model using the agentic, static-analysis-guided verification framework of SACTOR, which performs structure-aware, two-phase (unidiomatic to idiomatic) translation with foreign-function-interface (FFI)-based end-to-end (E2E) testing. We report success rate, idiomaticity (Clippy lint counts, unsafe-code fraction), and failure-mode analyses, and compare our fine-tuned model against baseline Qwen3-27B and other LLMs evaluated under the same framework.

cs.SE

Theory, Experience, and Instinct: A Glimpse Into AAA Game Processes and How UX Leaders Navigate Pre-Production

Foundational decisions shape a project's long-term trajectory, a dynamic that becomes especially evident in the inherent complexity of game pre-production. However, academic frameworks often see limited uptake at this stage, as they do not readily map onto industry contexts, production constraints, and cross-functional workflows. To better understand how design decisions are made in practice, we conducted interviews with 15 UX leaders from the AAA (triple-A) games industry. Our findings show that early UX decisions emerge from a dynamic blend of theory, experience, and intuition. In cross-functional structures (such as strike and competency teams), UX leaders collaboratively align player needs, technical feasibility, and creative vision. These decision-making processes involve translating academic concepts into production-ready insights, codifying experiential knowledge into reusable practices, and relying on informed intuition amid uncertainty. We argue that meaningful impact requires academia to develop malleable conceptual tools that integrate with practitioners' highly adaptive design processes. We conclude by discussing how existing frameworks might be adapted to connect academic insights with AAA workflows. Rather than prescriptive directives, we offer these as starting points for discussion that support strike and competency teams through shared language, reusable design systems, and strategies for collaborative, context-sensitive decision-making.

cs.HC

GRASP: Grounded CoT Reasoning with Dual-Stage Optimization for Multimodal Sarcasm Target Identification

Moving beyond the traditional binary classification paradigm of Multimodal Sarcasm Detection, Multimodal Sarcasm Target Identification (MSTI) presents a more formidable challenge, requiring precise localization of fine-grained targets such as textual phrases and visual regions. Existing approaches predominantly rely on implicit cross-modal alignment, offering limited interpretability and suboptimal fine-grained localization. To address these limitations, we propose GRASP, Grounded Chain-of-Thought ReAsoning with Dual-Stage Optimization for Multimodal Sarcasm Prediction and Target Identification, a framework that integrates visual grounding with explicit Chain-of-Thought (CoT) reasoning to move beyond black-box MSTI. Specifically, we curate MSTI-MAX, a refined dataset that mitigates class imbalance and enriches multimodal sarcasm cues. We introduce Grounded CoT reasoning, which explicitly anchors sarcasm-related visual regions within the reasoning trajectory and prompts the model to articulate rationales before predicting the final classification labels and sarcasm targets. Furthermore, we employ a dual-stage outcome-supervised joint optimization strategy: Supervised Fine-Tuning with a coordinate-aware weighted loss, followed by Fine-Grained Target Policy Optimization. Extensive experiments demonstrate that GRASP outperforms existing baselines in fine-grained sarcasm target identification across modalities, and an LLM-as-a-Judge evaluation quantitatively measures the quality of internal reasoning chains. Our dataset and source code will be released on GitHub.

cs.CL

Spectra of high-dimensional sparse random geometric graphs

We determine the limiting empirical spectral distribution of sparse high-dimensional random geometric graphs. The vertices are independent uniform points on the unit sphere $S^{d-1}$, and two vertices are joined when their inner product exceeds a threshold chosen to give edge density $p$. The edges therefore have the same marginal probabilities as in an Erdős--Rényi graph, but the latent geometry introduces dependence among them. We show that these correlations are asymptotically invisible to the global spectrum in two sparse regimes. If $p\to0$, $np\to\infty$, and $d=Ω(np\log(1/p))$, then the empirical spectral distribution of $A/\sqrt{np}$ converges in probability to the semicircle law. If $p=α/n$ for a fixed $α>0$ and $d=ω(\log n)$, then the empirical spectral distribution of $A/\sqrtα$ converges in probability to the limiting spectral distribution of $\mathcal G(n,α/n)$. The proof combines the moment method with a cluster expansion that decomposes geometric dependence into weak local interactions, allowing us to control every fixed walk pattern in the moment calculation.

math.PR

Human Cognition in Machines: A Unified Perspective of World Models

This report of world models distinguishes prior works by the cognitive functions they innovate. Many works claim an almost human-like cognitive capability in their world models. To evaluate these claims requires a proper grounding in first principles from human and machine cognition theory. In moving towards human-like world models we present a conceptual unified framework for world models that fully incorporates all the cognitive functions (i.e., memory, perception, language, reasoning, imagining, motivation, and metacognition) and identify gaps in existing research as a guide for future states of the art. In particular, we find that motivation (especially intrinsic motivation) and metacognition remain drastically under-researched, and we propose concrete directions to address these gaps informed by active inference and global workspace theory. We also introduce epistemic world models, a new category encompassing agent frameworks for scientific discovery that operate over structured knowledge. Our taxonomy, applied to video, embodied, and epistemic world models, suggests research directions where prior taxonomies have not.

cs.RO

CR-Seg: Attention-Guided and CoT-Enhanced Coarse-to-Refined Reasoning Segmentation

Reasoning segmentation aims to segment target objects described by complex language through joint visual-textual reasoning. Existing methods typically rely on either learned semantic tokens to bridge Multimodal Large Language Models (MLLMs) and segmentation models, suffering from difficult cross-modal alignment, or explicit spatial prompts such as bounding boxes, which may lose holistic response semantics. To address these limitations, we propose Attention-Guided and CoT-Enhanced Coarse-to-Refined Reasoning Segmentation, termed CR-Seg, a two-stage framework for coarse-to-refined reasoning segmentation. Specifically, we design an Extract Attention Maps and Points (EAP) module to extract attention maps for coarse target localization and select informative points, both of which are fed into SAM for mask refinement. To alleviate reasoning--answer inconsistency, we further introduce Global-to-Local Chain-of-Thought (GLCoT), which guides the model to reason progressively from global scene context to local target details. Extensive experiments on reasoning segmentation benchmarks demonstrate the effectiveness of CR-Seg.

cs.CV

Direct Observation of Chemical Short-Range Order in CoCrNi Alloy Using Neutron Diffraction

This study provides experimental evidence of chemical short-range order (CSRO) in the equiatomic CoCrNi alloy, identified through neutron diffraction. The phenomenon manifests as a distinct diffuse peak at Q = 1.85 A-1, the intensity increases under thermodynamically favorable conditions for CSRO development such as prolonged aging (100 h and 240 h) at 748 K or shorter aging (24 h) at slightly higher temperature (798 K). The degree of ordering was measured by integrating the diffuse scattering intensity, revealing that the gas-atomized sample, i.e. the sample with the least amount of CSRO, still displays approximately 70% of the CSRO level observed in the sample subsequently aged for 240 h at 748 K, i.e. the sample with the highest amount of CSRO produced in this study. Predictive atomistic simulations reproduced both the presence and position of the diffuse peak, while two-dimensional fast Fourier transform (FT-2D) analyses indicated that reflections at (1 1/2 0) within the <001> zone axis originate from some structural projections associated with like D022, Pt2Mo and D1a motifs. Complementary small-angle neutron scattering (SANS) measurements identified Ni-rich, disk-shaped domains with radii of approximately 11 A and thicknesses of about 1 A, consistent with nanoscale CSRO characteristic length scale. These findings demonstrate that CSRO is an intrinsic and energetically favorable feature of the CoCrNi system, remaining stable even under rapid solidification and further enhanced by low-temperature aging. Combined use of neutron diffraction and atomistic modeling provides a framework for probing local ordering phenomena in multi-principal element alloys (MPEAs).

cond-mat.mtrl-sci

TableTale: Reviving the Narrative Interplay Between Data Tables and Text in Scientific Papers

Data tables play a central role in scientific papers. However, their meaning is often co-constructed with surrounding text through narrative interplay, making comprehension cognitively demanding for readers. In this work, we explore how interfaces can better support this reading process. We conducted a formative study that revealed key characteristics of text-table narrative interplay, including linking mechanisms, multi-granularity alignments, and mention typologies, as well as a layered framework of readers' intents. Informed by these insights, we present TableTale, an augmented reading interface that enriches text with data tables at multiple granularities, including paragraphs, sentences, and mentions. TableTale automatically constructs a document-level linking schema within the paper and progressively renders cascade visual cues on text and tables that unfold as readers move through the text. A within-subject study with 24 participants showed that TableTale reduced cognitive workload and improved reading efficiency, demonstrating its potential to enhance paper reading and inform future reading interface design.

cs.HC

Nonequilibrium chemical short-range order in metallic alloys

Metallic alloys are routinely subjected to nonequilibrium processes during manufacturing, such as rapid solidification and thermomechanical processing. It has been suggested in the high-entropy alloy literature that chemical short-range order (SRO) could offer a new knob to tailor materials properties. While evidence of the effect of SRO on materials properties accumulates, the state of SRO evolution during alloy manufacturing remains obscure. Here, we employ high-fidelity atomistic simulations to track SRO evolution during the solidification and thermomechanical processing of alloys. Our investigation reveals that alloy processing can lead to nonequilibrium steady-states of SRO that are different from any equilibrium state. The mechanism behind nonequilibrium SRO formation is shown to be an inherent ordering bias present in nonequilibrium events. These results demonstrate that conventional manufacturing processes provide pathways for tuning SRO that lead to a broad nonequilibrium spectrum of SRO states beyond the equilibrium design space of alloys.

cond-mat.mtrl-sci

CodeAgents: A Token-Efficient Framework for Codified Multi-Agent Reasoning in LLMs

Effective prompt design is essential for improving the planning capabilities of large language model (LLM)-driven agents. However, existing structured prompting strategies are typically limited to single-agent, plan-only settings, and often evaluate performance solely based on task accuracy - overlooking critical factors such as token efficiency, modularity, and scalability in multi-agent environments. To address these limitations, we introduce CodeAgents, a prompting framework that codifies multi-agent reasoning and enables structured, token-efficient planning in multi-agent systems. In CodeAgents, all components of agent interaction - Task, Plan, Feedback, system roles, and external tool invocations - are codified into modular pseudocode enriched with control structures (e.g., loops, conditionals), boolean logic, and typed variables. This design transforms loosely connected agent plans into cohesive, interpretable, and verifiable multi-agent reasoning programs. We evaluate the proposed framework across three diverse benchmarks - GAIA, HotpotQA, and VirtualHome - using a range of representative LLMs. Results show consistent improvements in planning performance, with absolute gains of 3-36 percentage points over natural language prompting baselines. On VirtualHome, our method achieves a new state-of-the-art success rate of 56%. In addition, our approach reduces input and output token usage by 55-87% and 41-70%, respectively, underscoring the importance of token-aware evaluation metrics in the development of scalable multi-agent LLM systems. The code and resources are available at: https://anonymous.4open.science/r/CodifyingAgent-5A86

cs.AI

Machine learning potentials for modeling alloys across compositions

Materials properties depend strongly on chemical composition, i.e., the relative amounts of each chemical element. Changes in composition lead to entirely different chemical arrangements, which vary in complexity from perfectly ordered (i.e., stoichiometric compounds) to completely disordered (i.e., solid solutions). Accurately capturing this range of chemical arrangements remains a major challenge, limiting the predictive accuracy of machine learning potentials (MLPs) in materials modeling. Here, we combine information theory and machine learning to optimize the sampling of chemical motifs and design MLPs that effectively capture the behavior of metallic alloys across their entire compositional and structural landscape. The effectiveness of this approach is demonstrated by predicting the compositional dependence of various material properties - including stacking-fault energies, short-range order, heat capacities, and phase diagrams - for the AuPt and CuAu binary alloys, the ternary CrCoNi, and the TiTaVW high-entropy alloy. Extensive comparison against experimental data demonstrates the robustness of this approach in enabling materials modeling with high physical fidelity.

cond-mat.mtrl-sci

GPMFS: Global Foundation and Personalized Optimization for Multi-Label Feature Selection

As artificial intelligence methods are increasingly applied to complex task scenarios, high dimensional multi-label learning has emerged as a prominent research focus. At present, the curse of dimensionality remains one of the major bottlenecks in high-dimensional multi-label learning, which can be effectively addressed through multi-label feature selection methods. However, existing multi-label feature selection methods mostly focus on identifying global features shared across all labels, which overlooks personalized characteristics and specific requirements of individual labels. This global-only perspective may limit the ability to capture label-specific discriminative information, thereby affecting overall performance. In this paper, we propose a novel method called GPMFS (Global Foundation and Personalized Optimization for Multi-Label Feature Selection). GPMFS firstly identifies global features by exploiting label correlations, then adaptively supplements each label with a personalized subset of discriminative features using a threshold-controlled strategy. Experiments on multiple real-world datasets demonstrate that GPMFS achieves superior performance while maintaining strong interpretability and robustness. Furthermore, GPMFS provides insights into the label-specific strength across different multi-label datasets, thereby demonstrating the necessity and potential applicability of personalized feature selection approaches.

cs.LG

Chemical-motif characterization of short-range order with E(3)-equivariant graph neural networks

Crystalline materials have atomic-scale fluctuations in their chemical composition that modulate various mesoscale properties. Establishing chemistry-microstructure relationships in such materials requires proper characterization of these chemical fluctuations. Yet, current characterization approaches (e.g., Warren-Cowley parameters) make only partial use of the complete chemical and structural information contained in local chemical motifs. Here we introduce a framework based on E(3)-equivariant graph neural networks that is capable of completely identifying chemical motifs in arbitrary crystalline structures with any number of chemical elements. This approach naturally leads to a proper information-theoretic measure for quantifying chemical short-range order (SRO) in chemically complex materials, and a reduced - but complete - representation of the chemical space. Our framework enables the correlation of any per-atom property with their corresponding local chemical motif, thereby offering novel avenues to explore structure-property relationships in chemically-complex materials. Using the MoTaNbTi high-entropy alloy as a test system, we demonstrate the versatility of this approach by evaluating the lattice strain associated with each chemical motif, and computing the temperature dependence of chemical-fluctuations length scale.

cond-mat.mtrl-sci

NFTracer: Tracing NFT Impact Dynamics in Transaction-flow Substitutive Systems with Visual Analytics

Impact dynamics are crucial for estimating the growth patterns of NFT projects by tracking the diffusion and decay of their relative appeal among stakeholders. Machine learning methods for impact dynamics analysis are incomprehensible and rigid in terms of their interpretability and transparency, whilst stakeholders require interactive tools for informed decision-making. Nevertheless, developing such a tool is challenging due to the substantial, heterogeneous NFT transaction data and the requirements for flexible, customized interactions. To this end, we integrate intuitive visualizations to unveil the impact dynamics of NFT projects. We first conduct a formative study and summarize analysis criteria, including substitution mechanisms, impact attributes, and design requirements from stakeholders. Next, we propose the Minimal Substitution Model to simulate substitutive systems of NFT projects that can be feasibly represented as node-link graphs. Particularly, we utilize attribute-aware techniques to embed the project status and stakeholder behaviors in the layout design. Accordingly, we develop a multi-view visual analytics system, namely NFTracer, allowing interactive analysis of impact dynamics in NFT transactions. We demonstrate the informativeness, effectiveness, and usability of NFTracer by performing two case studies with domain experts and one user study with stakeholders. The studies suggest that NFT projects featuring a higher degree of similarity are more likely to substitute each other. The impact of NFT projects within substitutive systems is contingent upon the degree of stakeholders' influx and projects' freshness.

cs.CE

Capturing short-range order in high-entropy alloys with machine learning potentials

Chemical short-range order (SRO) affects the distribution of elements throughout the solid-solution phase of metallic alloys, thereby modifying the background against which microstructural evolution occurs. Investigating such chemistry-microstructure relationships requires atomistic models that act at the appropriate length scales while capturing the intricacies of chemical bonds leading to SRO. Here we consider various approaches for the construction of training data sets for machine learning potentials (MLPs) for CrCoNi and evaluate their performance in capturing SRO and its effects on materials quantities of relevance for mechanical properties, such as stacking-fault energy and phase stability. It is demonstrated that energy accuracy on test sets often does not correlate with accuracy in capturing material properties, which is fundamental in enabling large-scale atomistic simulations of metallic alloys with high physical fidelity. Based on this analysis we systematically derive design principles for the rational construction of MLPs that capture SRO in the crystal and liquid phases of alloys.

cond-mat.mtrl-sci

Quantifying chemical short-range order in metallic alloys

Metallic alloys often form phases - known as solid solutions - in which chemical elements are spread out on the same crystal lattice in an almost random manner. The tendency of certain chemical motifs to be more common than others is known as chemical short-range order (SRO) and it has received substantial consideration in alloys with multiple chemical elements present in large concentrations due to their extreme configurational complexity (e.g., high-entropy alloys). Short-range order renders solid solutions "slightly less random than completely random", which is a physically intuitive picture, but not easily quantifiable due to the sheer number of possible chemical motifs and their subtle spatial distribution on the lattice. Here we present a multiscale method to predict and quantify the SRO state of an alloy with atomic resolution, incorporating machine learning techniques to bridge the gap between electronic-structure calculations and the characteristic length scale of SRO. The result is an approach capable of predicting SRO length scale in agreement with experimental measurements while comprehensively correlating SRO with fundamental quantities such as local lattice distortions. This work advances the quantitative understanding of solid-solution phases, paving the way for SRO rigorous incorporation into predictive mechanical and thermodynamic models.

cond-mat.mtrl-sci

Zero-shot Compound Expression Recognition with Visual Language Model at the 6th ABAW Challenge

Conventional approaches to facial expression recognition primarily focus on the classification of six basic facial expressions. Nevertheless, real-world situations present a wider range of complex compound expressions that consist of combinations of these basics ones due to limited availability of comprehensive training datasets. The 6th Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW) offered unlabeled datasets containing compound expressions. In this study, we propose a zero-shot approach for recognizing compound expressions by leveraging a pretrained visual language model integrated with some traditional CNN networks.

cs.CV

Why Change My Design: Explaining Poorly Constructed Visualization Designs with Explorable Explanations

Although visualization tools are widely available and accessible, not everyone knows the best practices and guidelines for creating accurate and honest visual representations of data. Numerous books and articles have been written to expose the misleading potential of poorly constructed charts and teach people how to avoid being deceived by them or making their own mistakes. These readings use various rhetorical devices to explain the concepts to their readers. In our analysis of a collection of books, online materials, and a design workshop, we identified six common explanation methods. To assess the effectiveness of these methods, we conducted two crowdsourced studies (each with N = 125) to evaluate their ability to teach and persuade people to make design changes. In addition to these existing methods, we brought in the idea of Explorable Explanations, which allows readers to experiment with different chart settings and observe how the changes are reflected in the visualization. While we did not find significant differences across explanation methods, the results of our experiments indicate that, following the exposure to the explanations, the participants showed improved proficiency in identifying deceptive charts and were more receptive to proposed alterations of the visualization design. We discovered that participants were willing to accept more than 60% of the proposed adjustments in the persuasiveness assessment. Nevertheless, we found no significant differences among different explanation methods in convincing participants to accept the modifications.

cs.HC