SearcharxivSearch

arXiv subjects

Yiyang Luo

Publications and source records attributed to Yiyang Luo.

12 recordsLinked to original sources

ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance

LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence. In trader runs with DeepSeek V4 Pro and Gemini 3.5 Flash, visible rules reduce but do not eliminate rejected actions, and incentive or persona framing shifts behavior. A bridge study shows that trader rationales can mislead an independent monitor unless enforcement evidence is shown. In monitoring, simple structured baselines either match or exceed prompt-only LLMs. The results frame financial compliance evaluation as an audit of rule-grounded actions and evidence use, rather than a single compliance score.

cs.AI

Triplet-Block Diffusion RWKV

Causal Transformer language models suffer from strictly sequential decoding and a quadratic per-step attention cost. While linear-time causal models and discrete diffusion models each address these weaknesses, their integration remains inherently inconsistent: diffusion requires bidirectional attention, while causal models are unidirectional. To unify these architectures, we propose $B^3D-RWKV$, a diffusion RWKV variant that integrates the model's $O(L)$ inference efficiency with parallel, bidirectional discrete-diffusion through a \emph{triplet-block layout} method. $B^3D-RWKV-7.2B$ reaches comparable accuracy on an 8-task suite versus existing models while significantly outperforming baselines in decoding throughput with an average of $\mathbf{1.6\times}$ speedup.

cs.CL

SciDER: Scientific Data-centric End-to-end Researcher

While large language models accelerate scientific discovery, existing agents face severe limitations in adaptability, domain generalization, and multimodal scalability, often struggling to autonomously process raw, domain-specific experimental data. To overcome these barriers, we introduce SciDER, a multi-agent system designed to flexibly automate the entire research lifecycle. This framework employs a novel data-centric approach and integrates a dynamic multimodal skill system across four specialized sub-agents. Specifically, an ideation agent generates novel hypotheses via Evolutionary Idea Search, a data analysis agent systematically structures raw data, an experimentation agent synthesizes executable code grounded in dataset characteristics, and a critic agent drives iterative self-refinement. To democratize open-source scientific discovery, we release OpenSciDER-SFT-8K, a high-quality execution trajectory dataset, alongside the OpenSciDER-27B fine-tuned model. Across six benchmarks, SciDER and OpenSciDER obtain competitive or leading results, with especially strong gains on data-centric analysis, end-to-end research execution, and multimodal scientific visualization. By integrating data analysis with experimental execution, SciDER bridges the gap between abstract scientific reasoning and reproducible experimentation synthesis.

cs.AI

Geo-LLaVA: A Large Multi-Modal Model for Solving Geometry Math Problems with Meta In-Context Learning

Geometry mathematics problems pose significant challenges for large language models (LLMs) because they involve visual elements and spatial reasoning. Current methods primarily rely on symbolic character awareness to address these problems. Considering geometry problem solving is a relatively nascent field with limited suitable datasets and currently almost no work on solid geometry problem solving, we collect a geometry question-answer dataset by sourcing geometric data from Chinese high school education websites, referred to as GeoMath. It contains solid geometry questions and answers with accurate reasoning steps as compensation for existing plane geometry datasets. Additionally, we propose a Large Multi-modal Model (LMM) framework named Geo-LLaVA, which incorporates retrieval augmentation with supervised fine-tuning (SFT) in the training stage, called meta-training, and employs in-context learning (ICL) during inference to improve performance. Our fine-tuned model with ICL attains the state-of-the-art performance of 65.25% and 42.36% on selected questions of the GeoQA dataset and GeoMath dataset respectively with proper inference steps. Notably, our model initially endows the ability to solve solid geometry problems and supports the generation of reasonable solid geometry picture descriptions and problem-solving steps. Our research sets the stage for further exploration of LLMs in multi-modal math problem-solving, particularly in geometry math problems.

cs.CV

ViRED: Prediction of Visual Relations in Engineering Drawings

To accurately understand engineering drawings, it is essential to establish the correspondence between images and their description tables within the drawings. Existing document understanding methods predominantly focus on text as the main modality, which is not suitable for documents containing substantial image information. In the field of visual relation detection, the structure of the task inherently limits its capacity to assess relationships among all entity pairs in the drawings. To address this issue, we propose a vision-based relation detection model, named ViRED, to identify the associations between tables and circuits in electrical engineering drawings. Our model mainly consists of three parts: a vision encoder, an object encoder, and a relation decoder. We implement ViRED using PyTorch to evaluate its performance. To validate the efficacy of ViRED, we conduct a series of experiments. The experimental results indicate that, within the engineering drawing dataset, our approach attained an accuracy of 96\% in the task of relation prediction, marking a substantial improvement over existing methodologies. The results also show that ViRED can inference at a fast speed even when there are numerous objects in a single engineering drawing.

cs.CV

Context-Aware Indoor Point Cloud Object Generation through User Instructions

Indoor scene modification has emerged as a prominent area within computer vision, particularly for its applications in Augmented Reality (AR) and Virtual Reality (VR). Traditional methods often rely on pre-existing object databases and predetermined object positions, limiting their flexibility and adaptability to new scenarios. In response to this challenge, we present a novel end-to-end multi-modal deep neural network capable of generating point cloud objects seamlessly integrated with their surroundings, driven by textual instructions. Our model revolutionizes scene modification by enabling the creation of new environments with previously unseen object layouts, eliminating the need for pre-stored CAD models. Leveraging Point-E as our generative model, we introduce innovative techniques such as quantized position prediction and Top-K estimation to address the issue of false negatives resulting from ambiguous language descriptions. Furthermore, we conduct comprehensive evaluations to showcase the diversity of generated objects, the efficacy of textual instructions, and the quantitative metrics, affirming the realism and versatility of our model in generating indoor objects. To provide a holistic assessment, we incorporate visual grounding as an additional metric, ensuring the quality and coherence of the scenes produced by our model. Through these advancements, our approach not only advances the state-of-the-art in indoor scene modification but also lays the foundation for future innovations in immersive computing and digital environment creation.

cs.CV

Zero-shot Generative Linguistic Steganography

Generative linguistic steganography attempts to hide secret messages into covertext. Previous studies have generally focused on the statistical differences between the covertext and stegotext, however, ill-formed stegotext can readily be identified by humans. In this paper, we propose a novel zero-shot approach based on in-context learning for linguistic steganography to achieve better perceptual and statistical imperceptibility. We also design several new metrics and reproducible language evaluations to measure the imperceptibility of the stegotext. Our experimental results indicate that our method produces $1.926\times$ more innocent and intelligible stegotext than any other method.

cs.CL

Lost in Overlap: Exploring Logit-based Watermark Collision in LLMs

The proliferation of large language models (LLMs) in generating content raises concerns about text copyright. Watermarking methods, particularly logit-based approaches, embed imperceptible identifiers into text to address these challenges. However, the widespread usage of watermarking across diverse LLMs has led to an inevitable issue known as watermark collision during common tasks, such as paraphrasing or translation. In this paper, we introduce watermark collision as a novel and general philosophy for watermark attacks, aimed at enhancing attack performance on top of any other attacking methods. We also provide a comprehensive demonstration that watermark collision poses a threat to all logit-based watermark algorithms, impacting not only specific attack scenarios but also downstream applications.

cs.CL

Group-velocity-locked vector soliton molecules in a birefringence-enhanced fiber laser

Physics phenomena of multi-soliton complexes have enriched the life of dissipative solitons in fiber lasers. By developing a birefringence-enhanced fiber laser, we report the first experimental observation of group-velocity-locked vector soliton (GVLVS) molecules. The birefringence-enhanced fiber laser facilitates the generation of GVLVSs, where the two orthogonally polarized components are coupled together to form a multi-soliton complex. Moreover, the interaction of repulsive and attractive forces between multiple pulses binds the particle-like GVLVSs together in time domain to further form compound multi-soliton complexes, namely GVLVS molecules. By adopting the polarization-resolved measurement, we show that the two orthogonally polarized components of the GVLVS molecules are both soliton molecules supported by the strongly modulated spectral fringes and the double-humped intensity profiles. Additionally, GVLVS molecules with various soliton separations are also observed by adjusting the pump power and the polarization controller.

physics.optics

Group velocity locked vector dissipative solitons in a high repetition rate fiber laser

Vectorial nature of dissipative solitons (DSs) with high repetition rates is studied for the first time in a normal-dispersion fiber laser. Despite the fact that the formed DSs are strongly chirped and the repetition rate is greater than 100 MHz, polarization locked and polarization rotating group velocity locked vector DSs can be formed under 129.3 MHz fundamental mode-locking and 258.6 MHz harmonic mode-locking of the fiber laser, respectively. The two orthogonally polarized components of these vector DSs possess distinctly different central wavelengths and travel together at the same group velocity in the laser cavity, resulting in a gradual spectral edge and small steps on the optical spectra, which can be considered as an auxiliary indicator of the group velocity locked vector DSs.

physics.optics

Dynamics of dissipative solitons in a high repetition rate normal-dispersion erbium-doped fiber laser

The dynamics of dissipative solitons (DSs) are explored in a high repetition rate normal-dispersion erbium-doped fiber laser for the first time. Despite of the high fundamental repetition rate of 129 MHz and thus the low pulse energy, a DS train with a dechirped pulse width of 418 fs, period-doubling of single and dual DSs, as well as 258 MHz 2nd-order harmonic mode-locking of DSs can be observed in the fiber laser with increasing pump power and appropriate settings. A transmitted semiconductor saturable absorber and a wavelength division multiplexer/isolator/tap hybrid module are employed to simplify the laser configuration, thus not only increasing the repetition rate, but also enhancing the stability and robustness of the fiber laser due to the commercial availability of all the components.

physics.optics

Scalar - vector soliton fiber lasers

Rapid progress in passively mode-locked fiber lasers is currently driven by the recent discovery of vector feature of mode-locking pulses, namely, the group velocity-locked vector solitons, the phase locked vector solitons, and the high-order vector solitons. Those vector solitons are fundamentally different from the previously known scalar solitons. Here, we report a fiber laser where the mode-locked pulse evolves as a vector soliton in the strong birefringent segment and is transformed into a regular scalar soliton after the polarizer within the laser cavity. The existence of solutions in a polarization-dependent cavity comprising a periodic combination of two distinct nonlinear waves is novel and likely to be applicable to various other nonlinear systems. For very large local birefringence, our laser approaches the working regime of vector soliton lasers, while it approaches scalar soliton fiber lasers under the conditions of very small birefringence.

physics.optics