SearcharxivSearch

arXiv subjects

Yujia Li

Publications and source records attributed to Yujia Li.

At least 19 recordsLinked to original sources

HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes

Multimodal harmful meme detection is typically formulated as image--text harmfulness classification. A model may correctly predict harmfulness while misidentifying the attacked target or its supporting evidence. We therefore extend harmful meme detection with fine-grained target identification, asking what type of target is attacked, who is targeted, and where the target appears in the meme. The model predicts harmfulness for every meme and, for harmful memes, outputs the target category, target entity, textual mention, and visual region. To support this task, we introduce Meme3W, which unifies multiple public harmful meme datasets and provides human-verified annotations for harmful instances. We further introduce Joint Record Accuracy (JRA), a strict record-level metric requiring the harmfulness label and all target-identification fields to be jointly correct. Experiments with representative multimodal large language models reveal a substantial gap between harmfulness accuracy and JRA. To narrow this gap, we propose HarmTrace, an anchor-calibrated decoupled optimization framework. HarmTrace strengthens target-entity supervision through entity-aware supervised fine-tuning. It then applies Conditional Target-identification Policy Optimization (CTPO) to decouple harmfulness and target-identification advantages, restricting target-identification optimization to label-correct responses for harmful examples. CTPO uses a Virtual Positive Anchor (VPA) as a fully correct reference for target-identification advantage normalization. HarmTrace improves both JRA and harmfulness accuracy across the evaluated backbones, with JRA on the Qwen3-VL-8B backbone increasing from 17.58\% to 52.51\%. Our code is publicly available at https://github.com/llly1234/HarmTrace-for-Harmful-Memes.

cs.CV

Policy-Driven CT-Agent: Modeling Phase-Aware Diagnostic Control for Clinically Consistent CT Reasoning

Computed Tomography (CT) diagnosis often relies on dynamic selection of imaging phases, such as non-contrast, arterial, or venous phases, based on preliminary findings, clinical suspicion, and diagnostic guidelines. This phase-wise decision process is critical for reducing unnecessary radiation exposure while supporting timely staging and treatment planning. However, phase-selection protocols can vary across hospitals, regions, and guidelines, while most existing CT-based AI methods assume that all phases are available and focus on static tasks under a fixed imaging phase, failing to model whether additional phases are required. This limitation stems from heterogeneous multi-phase representations, the need for knowledge-guided phase control beyond visual cues, and the lack of supervision for phase-sufficiency decisions in existing datasets. To address these challenges, we propose Policy-Driven CT-Agent (PD-CTAgent) for clinically consistent CT phase selection and diagnostic reasoning. PD-CTAgent introduces a Clinical Structure Abstraction Module (CSAM) to harmonize heterogeneous CT phases into a unified, phase-aware evidence representation. Based on this representation, a Knowledge-Guided Diagnostic Control Model (KDCM) evaluates phase sufficiency and iteratively requests additional phases when necessary. The policy-driven agent design further allows PD-CTAgent to flexibly follow different institutional, regional, or guideline-specific diagnostic protocols. Together, PD-CTAgent bridges static CT analysis and real-world clinical workflows. Experiments on two public datasets, LIDC and MCT-LTDiag, and one private dataset demonstrate its effectiveness and clinical consistency. Code will be made public upon acceptance.

cs.LG

From Brewing to Resolution: Tracing the Internal Lifecycle of Code Reasoning in LLMs

Standard accuracy metrics cannot explain why LLMs handle variable tracking but fail on semantically equivalent loops. We study an internal lifecycle of code reasoning in which models first brew the answer, making it linearly recoverable many layers before it becomes self-decodable, and then diverge into one of four resolution outcomes: Resolved, Overprocessed, Misresolved, or Unresolved. Understanding this lifecycle matters because similar task accuracies can mask fundamentally different failure modes that surface-level evaluation cannot detect. We introduce a dual diagnostic framework pairing layer-wise linear probing with Context-Stripped Decoding (CSD) and apply it to six code-reasoning task families across 16 models spanning Qwen, Llama, and DeepSeek architectures. All four outcomes carry substantial mass in every task family: overall Resolved is only 41.5%, with multiple tasks below 30%. Controlled sweeps over structure, depth, and operators expose task-specific failure bottlenecks: Function Call Resolved plunges from 61.1% to 2.5% as call depth increases from one to three. Across architectures and scales, the brewing scaffold remains stable, with normalized brewing duration 24-42% across all 16 models, while resolution success varies with capability. This indicates that the scaffold is a stable empirical regularity across the tested decoder-only Transformer families, whereas resolution success covaries with capability, scale, and training. Code: https://github.com/euyis1019/llm-brewing

cs.AI

Beyond Text Following: Repairable Arbitration Reversals in Audio-Language Models

Audio-language models (ALMs) often follow text that conflicts with audio, even when the audio evidence is clear. This raises a basic question: is the audio-supported answer unavailable, or is it represented but overridden by the conflicting text? We examine this question using a same-audio counterfactual that keeps the audio fixed, removes only the conflicting text, and measures the resulting shift in model preference. Across five ALMs and four conflict tasks, 64.1% of conflict samples show a sign flip: the same-audio branch prefers the audio-supported answer, whereas the joint branch prefers the text-supported answer. This pattern suggests that the relevant audio evidence is encoded but loses in arbitration. Activation patching further localizes the reversal to answer-position computation, and patching effects closely track output candidate-score differences (Spearman rho=0.93). Using this diagnostic, we propose Gated Audio Counterfactual Logit Correction (GACL), a training-free decoding rule that interpolates between joint and same-audio scores. Under a strict 5 pp faithfulness-drop budget, GACL improves nAUC by 17.8 points over the best contrastive baseline and transfers without retuning to vision-text arbitration (up to +40.5 pp).

cs.SD

A Coupled V2G Equilibrium Model of Electric Vehicle and Power System Interactions

Vehicle-to-grid (V2G) technology empowers electric vehicles (EVs) to act as mobile energy resources, providing critical support to power systems, especially under stressed conditions. To understand the economic mechanism driving V2G participation and its benefits to power grid, this paper proposes a multi-player coupled equilibrium framework that models the bidirectional interactions between power grid operations and EV routing, incorporating charging and discharging choice in a preprocessed feasible path generation procedure. Energy prices are endogenously determined by market clearance conditions. We formulate the overall problem as a Variational Inequality that unite the decision-making of Distribution System Operator, Charging Network Operator, Load Serving Entities, and EV drivers. Numerical studies validate the framework under two stress scenarios: increased household load and power line outages. Results show that when EVs are incentivized by reduced generalized path costs, V2G is particularly effective in eliminating load shedding and reducing distribution locational marginal electricity prices. On the transportation side, V2G can lead to divergence in EV behavior between normal and scarcity conditions, and alter route choices yet improve overall trip economic.

eess.SY

Hybrid integrated narrow linewidth semiconductor laser based on the distributed feedback from an external deformed microcavity

Optical microcavities with rotational symmetry have been widely used for narrowing linewidth and reducing frequency noise, however, the narrow but wavelength dependent optical feedback restricts the narrow linewidth laser works only at some discrete wavelength matching the resonance of the microcavity. Here, we demonstrate a narrow linewidth semiconductor laser with continuous wavelength tunability by hybrid integrating a DFB laser chip with a deformed microcavity fabricated on a 220 nm SOI wafer. The deformed microcavity with vortex radius demonstrates the unique characteristics of unidirectional energy storage, wavelength self-adaptivity, and self-focusing of the Rayleigh scattering based distributed feedback. In addition, the strength of Rayleigh scattering is also significantly enhanced by the high numerical aperture silicon waveguide. The optical feedback signal measured by the optical frequency domain reflectometry (OFDR) shows that the deformed microcavity can effectively lengthen the equivalent propagation distance without wavelength dependence. With the wavelength self-adaptive optical feedback from the deformed microcavity, the intrinsic linewidth of a DFB laser diode is narrowed to 525 Hz and the side mode suppression ratio (SMSR) is improved to 76 dB in a maximum allowable continuous wavelength tuning range of 2.25 nm. The frequency noise and relative intensity noise (RIN) are reduced to 2.98 Hz2 /Hz and -148.74 dB/Hz at the offset frequency of 1 MHz, respectively. The work demonstrated here paves a new way for integrated tunable narrow linewidth lasers, which are of crucial importance in high-speed communication and high-precision spectroscopy

physics.optics

From Net Load Modifiers to Firm Capacity: The Role of Distributed Energy Resources in Resource Adequacy

Distributed energy resources (DERs) such as rooftop solar, batteries, demand response, and electric vehicles can contribute to power system reliability, yet their performance is difficult to translate into firm resource adequacy (RA) capacity across jurisdictions. Existing analyses often locate this difficulty within individual technical requirements, such as metering, accreditation, or dispatch performance, but give less attention to how constraints at one stage carry over to the next. This review traces the RA participation pathway through five stages: load forecasting, registration and classification, metering and verification, capacity accreditation, and performance obligations. We synthesize literature, tariffs, market manuals, and regulatory documents from California, PJM, ISO-NE, Great Britain, and Ireland, spanning U.S. capacity markets and European capacity remuneration mechanisms. Across these frameworks, similar barriers recur despite different procurement models and regulatory structures, indicating that participation is constrained by cross-stage design, not jurisdiction-specific rules alone. We identify three cross-stage couplings through which capacity value is lost between stages: mismatches between resource classification and operational obligations, weak links between verification evidence and accreditation, and temporal misalignment between planning forecasts and scarcity-hour performance. The central finding is that compliance architecture, not DER technology alone, is often the binding constraint on translating DER capability into firm RA contributions. This points to reforms that codify cross-stage information handoffs, tie accreditation to auditable verification evidence, and refresh capacity values as deployment changes system conditions. Rather than adjusting individual stages in isolation, RA reform should redesign the participation pathway end-to-end.

eess.SY

Time Window-Based Netload Range Cost Curves for Coordinated Transmission and Distribution Planning Under Uncertainty

Mechanisms to coordinate transmission and distribution planning should be regulatory compliant and keep the spheres of DSO and TSO decisions separate, without requiring disclosure of proprietary data or unrealistic computationally expensive T&D co-simulations. The concept of Netload Range Cost Curves (NRCC) has been recently proposed as simple non-invasive form of coordinating T&D investments under distribution netload uncertainty. This paper extends the NRCC concept to accommodate the temporal dimension of the T&D planning process. We propose to compute a hierarchy of certified temporal interface products that represent the different levels of flexibility that distribution networks can provide transmission grids with at the planning stage. The first product (P1) maps distribution investment into scenario-robust, per-window service envelopes within which any TSO service call (to modify load within specified bounds) is guaranteed distribution-network-feasible. The second product (P2) adds lexicographic rebound minimization, preserving P1-optimal service capacity while certifying post-service recovery under three governance variants with qualitatively distinct rebound-budget responses. In our numerical results, based on a real distribution feeder, we compare the performance of our proposed time-window-based flexibility products to an atemporal product (P0) that offers a static bound on the aggregate distribution grid netload across all time periods. Our results demonstrate the superiority of our proposed products in properly valuing the benefits of incremental investments in storage to allow for temporal flexibility.

eess.SY

GreenRFM: Learning a resource-efficient radiology vision-language foundation model via supervision-centric pre-training

Radiology foundation models (RFMs) have largely inherited the scale-first recipe of natural-image vision--language pre-training. This recipe is difficult to deploy in 3D radiology, where training corpora are smaller, reports vary across institutions, and receiving hospitals often need local adaptation under privacy and compute constraints. We ask whether routine radiology reports can instead be converted into auditable diagnostic supervision that shapes the image encoder, text encoder, aligned space, and local-adaptation procedure. We develop GreenRFM, a supervision-centric pre-training framework organized around four empirical principles: More distilled, Ubiquitous, Semantic-enforcing, and Task-aligning (MUST) supervision. These principles convert noisy reports into structured diagnostic signals and use them to learn discriminative unimodal encoders plus an aligned image--text space for diagnosis-centered multimodal use. GreenRFM requires 24 GPU-hours on a single 24GB GPU (lightweight variant: 6GB VRAM, 4~hours) and reaches a zero-shot CT-RATE AUC of 84.8. Evaluations using more than 200,000 volumes from six institutions and two modalities show transfer to private clinical cohorts and to musculoskeletal MRI. On a local institutional cohort, computationally feasible retraining raises macro-AUC from 70.5 to 82.1. The aligned space also improves hepatocellular-carcinoma microvascular-invasion prediction and trans-arterial chemoembolization response analysis over established clinical scores. These results support supervision-centric pre-training as a practical route to resource-efficient, locally adaptable, diagnosis-centered radiology vision--language representations.

cs.CV

Tower of Babel in Cross-Cultural Communication: A Case Study of #Give Me a Chinese Name# Dialogues During the "TikTok Refugees'' Event

The sudden influx of "TikTok refugees'' into the Chinese platform RedNote in early 2025 created an unprecedented, large-scale online cross-cultural communication event between the West and East. Although prior HCI research has studied user behavior in social media, most work remains confined to monolingual or single-cultural contexts, leaving cross-linguistic and cultural dynamics underexplored. To address this gap, we focused on a particularly challenging cross-cultural encoding-decoding task that remains stubbornly beyond the reach of machine translation, i.e., foreign newcomers asking Chinese users for Chinese names, and examined how people collectively constructed a digital "Babel Tower'' through various information encoding strategies. We collected and analyzed over 70,000 comments from RedNote with a creative human-in-the-loop approach using large language models, deriving a systematic framework summarizing cross-cultural information encoding strategies, how they are combined and layered to complicate decoding, and how they relate to engagement metrics such as the number of likes.

cs.HC

Broadband tunable narrow-linewidth laser based on scattering-enhanced fiber covering E-S-C-L bands

This work demonstrates a broadband tunable narrow-linewidth laser based on scattering-enhanced fiber, covering the E-S-C-L wavelength bands from 1337.47 nm to 1631.39 nm, with a total tuning span of 293.92 nm. The laser employs two semiconductor optical amplifiers (SOAs) centered at 1420 nm and 1550 nm, which are connected into a single ring resonator via polarization multiplexing. Wavelength selection and tunability is realized using an ultra-broadband tunable filter based on a blazed grating. To suppress side longitude modes, an 18-meter-long femtosecond-laser-empowered random scattering fiber is utilized inside the cavity as a feedback medium, yielding an output linewidths between 1.54 kHz and 2.61 kHz. Benefited from the fast response of the galvanometer mirror and short relaxation time of SOAs, wavelength switching time is less than 1 ms under different tuning channels among the wavelength range of near 300 nm. The stable single-longitude-mode operation is maintained across the entire tuning range. The exceptionally broad tuning range and high spectral purity of the laser endow it with significant application potentials across a wide range of fields.

physics.optics

Hybrid integrated narrow linewidth laser with external distributed optical feedback from a silicon strip waveguide

External optical feedback via Rayleigh scattering from an integrated microresonator or an optical fiber has been demonstrated to significantly narrow the intrinsic linewidth of semiconductor lasers. Wavelength matching between the lasing cavity and the external high-Q microresonator is required to accumulate Rayleigh scattering based optical feedback. Optical fiber can provide Rayleigh scattering based optical feedback for any lasing wavelength. However, optical fibers hundreds of meters or even kilometers long are required for the accumulation of Rayleigh scattering based optical feedback, hindering the integration of narrow linewidth lasers. Here, we present an integrated scheme that collects distributed feedback signal with weak wavelength dependence by exploiting surface radiation in a silicon waveguide. The effects of waveguide width on the intensities of the surface radiation and distributed optical feedback signal are first numerically analyzed by introducing a collection coefficient. Numerical calculations show that a 1 {\mu}m-wide strip waveguide yields optimal performance for excitation and collection of distributed optical feedback, which is also experimentally verified by measuring the feedback signal with an optical frequency-domain reflectometry. Benefitting from the enhanced distributed optical feedback that is 34.72 dB higher than that in a single-mode fiber, the hybrid integrated laser demonstrates an intrinsic linewidth of 1.52 kHz, a side-mode suppression ratio (SMSR) of 74.71 dB, and a frequency noise of 24.44 Hz2/Hz. Furthermore, within a maximum allowable wavelength tuning range of 2.342 nm, the linewidth narrowing ratio depends little on the wavelength for all the waveguides with different widths.

physics.optics

Adaptive Federated Learning to Optimize Integrated Flows in Cyber-Physical Data Centers

Data centers play an increasingly critical role in societal digitalization, yet their rapidly growing energy demand poses significant challenges for sustainable operation. To enhance the energy efficiency of geographically distributed data centers, this paper formulates a multi-period optimization model that captures the interdependence of electricity, heat, and data flows. The optimization of such integrated multi-domain flows inherently involves mixed-integer formulations and the access to proprietary or sensitive datasets, which correspondingly exacerbate computational complexity and raise data-privacy concerns. To address these challenges, an adaptive federated learning-to-optimization approach is proposed, accounting for the heterogeneity of datasets across distributed data centers. To safeguard privacy, cryptography techniques are leveraged in both the learning and optimization processes. A model acceptance criterion with convergence guarantee is developed to improve learning performance and filter out potentially contaminated data, while a verifiable double aggregation mechanism is further proposed to simultaneously ensure privacy and integrity of shared data during optimization. Theoretical analysis and numerical simulations demonstrate that the proposed approach preserves the privacy and integrity of shared data, achieves near-optimal performance, and exhibits high computational efficiency, making it suitable for large-scale data center optimization under privacy constraints.

eess.SY

NeRF-based CBCT Reconstruction needs Normalization and Initialization

Cone Beam Computed Tomography (CBCT) is widely used in medical imaging. However, the limited number and intensity of X-ray projections make reconstruction an ill-posed problem with severe artifacts. NeRF-based methods have achieved great success in this task. However, they suffer from a local-global training mismatch between their two key components: the hash encoder and the neural network. Specifically, in each training step, only a subset of the hash encoder's parameters is used (local sparse), whereas all parameters in the neural network participate (global dense). Consequently, hash features generated in each step are highly misaligned, as they come from different subsets of the hash encoder. These misalignments from different training steps are then fed into the neural network, causing repeated inconsistent global updates in training, which leads to unstable training, slower convergence, and degraded reconstruction quality. Aiming to alleviate the impact of this local-global optimization mismatch, we introduce a Normalized Hash Encoder, which enhances feature consistency and mitigates the mismatch. Additionally, we propose a Mapping Consistency Initialization(MCI) strategy that initializes the neural network before training by leveraging the global mapping property from a well-trained model. The initialized neural network exhibits improved stability during early training, enabling faster convergence and enhanced reconstruction performance. Our method is simple yet effective, requiring only a few lines of code while substantially improving training efficiency on 128 CT cases collected from 4 different datasets, covering 7 distinct anatomical regions.

eess.IV

MS-Glance: Bio-Insipred Non-semantic Context Vectors and their Applications in Supervising Image Reconstruction

Non-semantic context information is crucial for visual recognition, as the human visual perception system first uses global statistics to process scenes rapidly before identifying specific objects. However, while semantic information is increasingly incorporated into computer vision tasks such as image reconstruction, non-semantic information, such as global spatial structures, is often overlooked. To bridge the gap, we propose a biologically informed non-semantic context descriptor, \textbf{MS-Glance}, along with the Glance Index Measure for comparing two images. A Global Glance vector is formulated by randomly retrieving pixels based on a perception-driven rule from an image to form a vector representing non-semantic global context, while a local Glance vector is a flattened local image window, mimicking a zoom-in observation. The Glance Index is defined as the inner product of two standardized sets of Glance vectors. We evaluate the effectiveness of incorporating Glance supervision in two reconstruction tasks: image fitting with implicit neural representation (INR) and undersampled MRI reconstruction. Extensive experimental results show that MS-Glance outperforms existing image restoration losses across both natural and medical images. The code is available at \url{https://github.com/Z7Gao/MSGlance}.

eess.IV

Longitudinal Causal Image Synthesis

Clinical decision-making relies heavily on causal reasoning and longitudinal analysis. For example, for a patient with Alzheimer's disease (AD), how will the brain grey matter atrophy in a year if intervened on the A-beta level in cerebrospinal fluid? The answer is fundamental to diagnosis and follow-up treatment. However, this kind of inquiry involves counterfactual medical images which can not be acquired by instrumental or correlation-based image synthesis models. Yet, such queries require counterfactual medical images, not obtainable through standard image synthesis models. Hence, a causal longitudinal image synthesis (CLIS) method, enabling the synthesis of such images, is highly valuable. However, building a CLIS model confronts three primary yet unmet challenges: mismatched dimensionality between high-dimensional images and low-dimensional tabular variables, inconsistent collection intervals of follow-up data, and inadequate causal modeling capability of existing causal graph methods for image data. In this paper, we established a tabular-visual causal graph (TVCG) for CLIS overcoming these challenges through a novel integration of generative imaging, continuous-time modeling, and structural causal models combined with a neural network. We train our CLIS based on the ADNI dataset and evaluate it on two other AD datasets, which illustrate the outstanding yet controllable quality of the synthesized images and the contributions of synthesized MRI to the characterization of AD progression, substantiating the reliability and utility in clinics.

eess.IV

Pump-locked microcavity Brillouin laser

Microcavity-based microlasers are the kernel light sources for integrating photonics and optoelectronics. The traditional pump light frequency locking mainly utilizes a complex system with optoelectronic feedback, which requires a high-cost narrow-linewidth pump laser and limits the application of microlasers in integrated optoelectronic systems. We propose to utilize Rayleigh scattering of microcavities to lock the frequency of the pump laser to the resonant frequency of the laser microcavity with an all-optical method. While compressing the linewidth of the pump laser, it can greatly improve the long-term stability of the optically pumped microcavity laser. In the experiment, the linewidth of the semiconductor pump laser is compressed from the MHz level to the kHz level. The microcavity Brillouin laser achieves an ultra-narrow intrinsic linewidth of 100 Hz, with an ultra-low frequency noise of 35 Hz2/Hz. The constructed microlaser obtains a locking time up to 1 hour, which does not require any temperature control or vibration isolation of the laser system. This work is the first demonstration to achieve an optically pump-locked microcavity Brillouin laser, which provides a stable and reliable low-cost experimental platform for ultra-narrow linewidth lasers, precision laser sensors, microwave-photonic signal synthesizer, and optomechanical systems.

physics.optics

Large Language Models as Analogical Reasoners

Chain-of-thought (CoT) prompting for language models demonstrates impressive performance across reasoning tasks, but typically needs labeled exemplars of the reasoning process. In this work, we introduce a new prompting approach, analogical prompting, designed to automatically guide the reasoning process of large language models. Inspired by analogical reasoning, a cognitive process in which humans draw from relevant past experiences to tackle new problems, our approach prompts language models to self-generate relevant exemplars or knowledge in the context, before proceeding to solve the given problem. This method presents several advantages: it obviates the need for labeling or retrieving exemplars, offering generality and convenience; it can also tailor the generated exemplars and knowledge to each problem, offering adaptability. Experimental results show that our approach outperforms 0-shot CoT and manual few-shot CoT in a variety of reasoning tasks, including math problem solving in GSM8K and MATH, code generation in Codeforces, and other reasoning tasks in BIG-Bench.

cs.LG