Searcharxiv⌕ Search

arXiv subjects

Yu Guo

Publications and source records attributed to Yu Guo.

At least 91 records · Page 5Linked to original sources

A Survey on Large Language Model-Based Social Agents in Game-Theoretic Scenarios

Game-theoretic scenarios have become pivotal in evaluating the social intelligence of Large Language Model (LLM)-based social agents. While numerous studies have explored these agents in such settings, there is a lack of a comprehensive survey summarizing the current progress. To address this gap, we systematically review existing research on LLM-based social agents within game-theoretic scenarios. Our survey organizes the findings into three core components: Game Framework, Social Agent, and Evaluation Protocol. The game framework encompasses diverse game scenarios, ranging from choice-focusing to communication-focusing games. The social agent part explores agents' preferences, beliefs, and reasoning abilities, as well as their interactions and synergistic effects on decision-making. The evaluation protocol covers both game-agnostic and game-specific metrics for assessing agent performance. Additionally, we analyze the performance of current social agents across various game scenarios. By reflecting on the current research and identifying future research directions, this survey provides insights to advance the development and evaluation of social agents in game-theoretic scenarios.

cs.CL↗

Boundary Learning by Using Weighted Propagation in Convolution Network

In material science, image segmentation is of great significance for quantitative analysis of microstructures. Here, we propose a novel Weighted Propagation Convolution Neural Network based on U-Net (WPU-Net) to detect boundary in poly-crystalline microscopic images. We introduce spatial consistency into network to eliminate the defects in raw microscopic image. And we customize adaptive boundary weight for each pixel in each grain, so that it leads the network to preserve grain's geometric and topological characteristics. Moreover, we provide our dataset with the goal of advancing the development of image processing in materials science. Experiments demonstrate that the proposed method achieves promising performance in both of objective and subjective assessment. In boundary detection task, it reduces the error rate by 7\%, which outperforms state-of-the-art methods by a large margin.

cs.CV↗

ViTaL: A Multimodality Dataset and Benchmark for Multi-pathological Ovarian Tumor Recognition

Ovarian tumor, as a common gynecological disease, can rapidly deteriorate into serious health crises when undetected early, thus posing significant threats to the health of women. Deep neural networks have the potential to identify ovarian tumors, thereby reducing mortality rates, but limited public datasets hinder its progress. To address this gap, we introduce a vital ovarian tumor pathological recognition dataset called \textbf{ViTaL} that contains \textbf{V}isual, \textbf{T}abular and \textbf{L}inguistic modality data of 496 patients across six pathological categories. The ViTaL dataset comprises three subsets corresponding to different patient data modalities: visual data from 2216 two-dimensional ultrasound images, tabular data from medical examinations of 496 patients, and linguistic data from ultrasound reports of 496 patients. It is insufficient to merely distinguish between benign and malignant ovarian tumors in clinical practice. To enable multi-pathology classification of ovarian tumor, we propose a ViTaL-Net based on the Triplet Hierarchical Offset Attention Mechanism (THOAM) to minimize the loss incurred during feature fusion of multi-modal data. This mechanism could effectively enhance the relevance and complementarity between information from different modalities. ViTaL-Net serves as a benchmark for the task of multi-pathology, multi-modality classification of ovarian tumors. In our comprehensive experiments, the proposed method exhibited satisfactory performance, achieving accuracies exceeding 90\% on the two most common pathological types of ovarian tumor and an overall performance of 85\%. Our dataset and code are available at https://github.com/GGbond-study/vitalnet.

eess.IV↗

Spherical Phase Metalenses: Intrinsic Suppression of Spherical Aberration via Equiphase Surface Modulation

Recent progress in large-scale metasurfaces requires phase profiles beyond traditional hyperbolic designs. We show hyperbolic phase distributions cause spherical aberration from mismatched light propagation geometry and unrealistic phase assumptions. By analyzing metalens fundamentals via isophase surfaces, we develop a spherical phase profile based on spherical wavefront theory. This method prevents spherical aberration, essential for wide-aperture metalenses. Simulations prove superior focusing: spherical phase reduces FWHM by 7.3% and increases peak intensity by 20.4% versus hyperbolic designs at 31.46 micron radius. Spherical phase maintains consistent focusing across radii, while hyperbolic phase shows strong correlation (R squared = 0.95) with aberration. We also propose a normal vector tracing metric to measure design aberrations. This work establishes a scalable framework for diffraction-limited metalenses.

physics.optics↗

BuildingBRep-11K: Precise Multi-Storey B-Rep Building Solids with Rich Layout Metadata

With the rise of artificial intelligence, the automatic generation of building-scale 3-D objects has become an active research topic, yet training such models still demands large, clean and richly annotated datasets. We introduce BuildingBRep-11K, a collection of 11 978 multi-storey (2-10 floors) buildings (about 10 GB) produced by a shape-grammar-driven pipeline that encodes established building-design principles. Every sample consists of a geometrically exact B-rep solid-covering floors, walls, slabs and rule-based openings-together with a fast-loading .npy metadata file that records detailed per-floor parameters. The generator incorporates constraints on spatial scale, daylight optimisation and interior layout, and the resulting objects pass multi-stage filters that remove Boolean failures, undersized rooms and extreme aspect ratios, ensuring compliance with architectural standards. To verify the dataset's learnability we trained two lightweight PointNet baselines. (i) Multi-attribute regression. A single encoder predicts storey count, total rooms, per-storey vector and mean room area from a 4 000-point cloud. On 100 unseen buildings it attains 0.37-storey MAE (87 \% within $\pm1$), 5.7-room MAE, and 3.2 m$^2$ MAE on mean area. (ii) Defect detection. With the same backbone we classify GOOD versus DEFECT; on a balanced 100-model set the network reaches 54 \% accuracy, recalling 82 \% of true defects at 53 \% precision (41 TP, 9 FN, 37 FP, 13 TN). These pilots show that BuildingBRep-11K is learnable yet non-trivial for both geometric regression and topological quality assessment

cs.LG↗

TMGBench: A Systematic Game Benchmark for Evaluating Strategic Reasoning Abilities of LLMs

The rapid advancement of large language models has accelerated their application in reasoning, with strategic reasoning drawing increasing attention. To evaluate the strategic reasoning capabilities of LLMs, game theory, with its concise structure, has become the preferred approach for many researchers. However, current research typically focuses on a limited selection of games, resulting in low coverage of game types. Additionally, classic game scenarios carry risks of data leakage, and the benchmarks used often lack extensibility, rendering them inadequate for evaluating state-of-the-art models. To address these challenges, we propose TMGBench, characterized by comprehensive game type coverage, diverse scenarios and flexible game organization. Specifically, we incorporate all 144 game types summarized by the Robinson-Goforth topology of 2x2 games, constructed as classic games in our benchmark; we also synthetize diverse, higher-quality game scenarios for each classic game, which we refer to as story-based games. Lastly, to provide a sustainable evaluation framework adaptable to increasingly powerful LLMs, we treat the aforementioned games as atomic units and organize them into more complex forms through sequential, parallel, and nested structures. We conducted a comprehensive evaluation of mainstream LLMs, covering tests on rational reasoning, reasoning robustness, Theory-of-Mind capabilities, and reasoning in complex game forms. The results revealed LLMs still have flaws in the accuracy and consistency of strategic reasoning processes, and their levels of mastery over Theory-of-Mind also vary. Additionally, SOTA models like o3-mini, Qwen3 and deepseek-reasoner, were also evaluated across the sequential, parallel, and nested game structures while the results highlighted the challenges posed by TMGBench.

cs.AI↗

Expanding Zero-Shot Object Counting with Rich Prompts

Expanding pre-trained zero-shot counting models to handle unseen categories requires more than simply adding new prompts, as this approach does not achieve the necessary alignment between text and visual features for accurate counting. We introduce RichCount, the first framework to address these limitations, employing a two-stage training strategy that enhances text encoding and strengthens the model's association with objects in images. RichCount improves zero-shot counting for unseen categories through two key objectives: (1) enriching text features with a feed-forward network and adapter trained on text-image similarity, thereby creating robust, aligned representations; and (2) applying this refined encoder to counting tasks, enabling effective generalization across diverse prompts and complex images. In this manner, RichCount goes beyond simple prompt expansion to establish meaningful feature alignment that supports accurate counting across novel categories. Extensive experiments on three benchmark datasets demonstrate the effectiveness of RichCount, achieving state-of-the-art performance in zero-shot counting and significantly enhancing generalization to unseen categories in open-world scenarios.

cs.CV↗

Instruct2See: Learning to Remove Any Obstructions Across Distributions

Images are often obstructed by various obstacles due to capture limitations, hindering the observation of objects of interest. Most existing methods address occlusions from specific elements like fences or raindrops, but are constrained by the wide range of real-world obstructions, making comprehensive data collection impractical. To overcome these challenges, we propose Instruct2See, a novel zero-shot framework capable of handling both seen and unseen obstacles. The core idea of our approach is to unify obstruction removal by treating it as a soft-hard mask restoration problem, where any obstruction can be represented using multi-modal prompts, such as visual semantics and textual instructions, processed through a cross-attention unit to enhance contextual understanding and improve mode control. Additionally, a tunable mask adapter allows for dynamic soft masking, enabling real-time adjustment of inaccurate masks. Extensive experiments on both in-distribution and out-of-distribution obstacles show that Instruct2See consistently achieves strong performance and generalization in obstruction removal, regardless of whether the obstacles were present during the training phase. Code and dataset are available at https://jhscut.github.io/Instruct2See.

cs.CV↗

SQLForge: Synthesizing Reliable and Diverse Data to Enhance Text-to-SQL Reasoning in LLMs

Large Language models (LLMs) have demonstrated significant potential in text-to-SQL reasoning tasks, yet a substantial performance gap persists between existing open-source models and their closed-source counterparts. In this paper, we introduce SQLForge, a novel approach for synthesizing reliable and diverse data to enhance text-to-SQL reasoning in LLMs. We improve data reliability through SQL syntax constraints and SQL-to-question reverse translation, ensuring data logic at both structural and semantic levels. We also propose an SQL template enrichment and iterative data domain exploration mechanism to boost data diversity. Building on the augmented data, we fine-tune a variety of open-source models with different architectures and parameter sizes, resulting in a family of models termed SQLForge-LM. SQLForge-LM achieves the state-of-the-art performance on the widely recognized Spider and BIRD benchmarks among the open-source models. Specifically, SQLForge-LM achieves EX accuracy of 85.7% on Spider Dev and 59.8% on BIRD Dev, significantly narrowing the performance gap with closed-source methods.

cs.CL↗

Observation of Genuine High-dimensional Multi-partite Non-locality in Entangled Photon States

Quantum information science has leaped forward with the exploration of high-dimensional quantum systems, offering greater potential than traditional qubits in quantum communication and quantum computing. To advance the field of high-dimensional quantum technology, a significant effort is underway to progressively enhance the entanglement dimension between two particles. An alternative effective strategy involves not only increasing the dimensionality but also expanding the number of particles that are entangled. We present an experimental study demonstrating multi-partite quantum non-locality beyond qubit constraints, thus moving into the realm of strongly entangled high-dimensional multi-particle quantum systems. In the experiment, quantum states were encoded in the path degree of freedom (DoF) and controlled via polarization, enabling efficient operations in a two-dimensional plane to prepare three- and four-particle Greenberger-Horne-Zeilinger (GHZ) states in three-level systems. Our experimental results reveal ways in which high-dimensional systems can surpass qubits in terms of violating local-hidden-variable theories. Our realization of multiple complex and high-quality entanglement technologies is an important primary step for more complex quantum computing and communication protocols.

quant-ph↗

Surpassing the Global Heisenberg Limit Using a High-effciency Quantum Switch

Indefinite causal orders have been shown to enable a precision of inverse square N in quantum parameter estimation, where N is the number of independent processes probed in an experiment. This surpasses the widely accepted ultimate quantum precision of the Heisenberg limit, 1/N. While a recent laboratory demonstration highlighted this phenomenon, its validity relies on postselection for it only accounted for a subset of the resources used. Achieving a true violation of the Heisenberg limit-considering photon loss, detection ineffciency, and other imperfections-remains an open challenge. Here, we present an ultrahigh-effciency quantum switch to estimate the geometric phase associated with a pair of conjugate position and momentum displacements embedded in a superposition of causal orders. Our experimental data demonstrate precision surpassing the global Heisenberg limit without artificially correcting for losses or imperfections. This work paves the way for quantum metrology advantages under more general and realistic constraints.

quant-ph↗

Degeneracy-Locked Optical Parametric Oscillator

Optical parametric oscillators (OPOs) are widely utilized in photonics as classical and quantum light sources. Conventional OPOs produce co-propagating signal and idler waves that can be either degenerately or non-degenerately phase-matched. This configuration, however, makes it challenging to separate signal and idler waves and also renders their frequencies highly sensitive to external disturbances. Here, we demonstrate a degeneracy-locked OPO achieved through backward phase matching in a submicron periodically-poled thin-film lithium niobate microresonator. While the backward phase matching establishes frequency degeneracy of the signal and idler, the backscattering in the waveguide further ensures phase-locking between them. Their interplay permits the locking of the OPO's degeneracy over a broad parameter space, resulting in deterministic degenerate OPO initiation and robust operation against both pump detuning and temperature fluctuations. This work thus provides a new approach for synchronized operations in nonlinear photonics and extends the functionality of optical parametric oscillators. With its potential for large-scale integration, it provides a chip-based platform for advanced applications, such as squeezed light generation, coherent optical computing, and investigations of complex nonlinear phenomena.

physics.optics↗

Quaternion Nuclear Norms Over Frobenius Norms Minimization for Robust Matrix Completion

Recovering hidden structures from incomplete or noisy data remains a pervasive challenge across many fields, particularly where multi-dimensional data representation is essential. Quaternion matrices, with their ability to naturally model multi-dimensional data, offer a promising framework for this problem. This paper introduces the quaternion nuclear norm over the Frobenius norm (QNOF) as a novel nonconvex approximation for the rank of quaternion matrices. QNOF is parameter-free and scale-invariant. Utilizing quaternion singular value decomposition, we prove that solving the QNOF can be simplified to solving the singular value $L_1/L_2$ problem. Additionally, we extend the QNOF to robust quaternion matrix completion, employing the alternating direction multiplier method to derive solutions that guarantee weak convergence under mild conditions. Extensive numerical experiments validate the proposed model's superiority, consistently outperforming state-of-the-art quaternion methods.

cs.CV↗

ePBR: Extended PBR Materials in Image Synthesis

Realistic indoor or outdoor image synthesis is a core challenge in computer vision and graphics. The learning-based approach is easy to use but lacks physical consistency, while traditional Physically Based Rendering (PBR) offers high realism but is computationally expensive. Intrinsic image representation offers a well-balanced trade-off, decomposing images into fundamental components (intrinsic channels) such as geometry, materials, and illumination for controllable synthesis. However, existing PBR materials struggle with complex surface models, particularly high-specular and transparent surfaces. In this work, we extend intrinsic image representations to incorporate both reflection and transmission properties, enabling the synthesis of transparent materials such as glass and windows. We propose an explicit intrinsic compositing framework that provides deterministic, interpretable image synthesis. With the Extended PBR (ePBR) Materials, we can effectively edit the materials with precise controls.

cs.GR↗

Practical Advantage of Classical Communication in Entanglement Detection

Entanglement is the cornerstone of quantum communication, yet conventional detection relies solely on local measurements. In this work, we present a unified theoretical and experimental framework demonstrating that one-way local operations and classical communication (1-LOCC) can significantly outperform purely local measurements in detecting high-dimensional quantum entanglement. By casting the entanglement detection problem as a semidefinite program (SDP), we derive protocols that minimize false negatives at fixed false-positive rates. A variational generative machine-learning algorithm efficiently searches over high-dimensional parameter spaces, identifying states and measurement strategies that exhibit a clear 1-LOCC advantage. Experimentally, we realize a genuine event-ready protocol on a three-dimensional photonic entanglement source, employing fiber delays as short-lived quantum memories. We implement rapid, FPGA-based sampling of the optimized probabilistic instructions, allowing Bob's measurement settings to adapt to Alice's outcomes in real time. Our results validate the predicted 1-LOCC advantage in a realistic noisy setting and reduce the experimental trials needed to certify entanglement. These findings mark a step toward scalable, adaptive entanglement detection methods crucial for quantum networks and computing, paving the way for more efficient generation and verification of high-dimensional entangled states.

quant-ph↗

In vivo dynamic optical coherence tomography of human skin with hardware- and software-based motion correction

In vivo application of dynamic optical coherence tomography (DOCT) is hindered by bulk motion of the sample. We demonstrate DOCT imaging of \invivo human skin by adopting a sample-fixation attachment to suppress bulk motion and a subsequent software motion correction to further reduce the effect of sample motion. The performance of the motion-correction method was assessed by DOCT image observation, statistical analysis of the mean DOCT values, and subjective image grading. Both the mean DOCT value analysis and subjective grading showed statistically significant improvement of the DOCT image quality. In addition, a previously unobserved high DOCT layer was identified though image observation, which may represent the stratum basale with high keratinocyte proliferation.

physics.optics↗

Seeing A 3D World in A Grain of Sand

We present a snapshot imaging technique for recovering 3D surrounding views of miniature scenes. Due to their intricacy, miniature scenes with objects sized in millimeters are difficult to reconstruct, yet miniatures are common in life and their 3D digitalization is desirable. We design a catadioptric imaging system with a single camera and eight pairs of planar mirrors for snapshot 3D reconstruction from a dollhouse perspective. We place paired mirrors on nested pyramid surfaces for capturing surrounding multi-view images in a single shot. Our mirror design is customizable based on the size of the scene for optimized view coverage. We use the 3D Gaussian Splatting (3DGS) representation for scene reconstruction and novel view synthesis. We overcome the challenge posed by our sparse view input by integrating visual hull-derived depth constraint. Our method demonstrates state-of-the-art performance on a variety of synthetic and real miniature scenes.

cs.CV↗

Training with Differential Privacy: A Gradient-Preserving Noise Reduction Approach with Provable Security

Deep learning models have been extensively adopted in various regions due to their ability to represent hierarchical features, which highly rely on the training set and procedures. Thus, protecting the training process and deep learning algorithms is paramount in privacy preservation. Although Differential Privacy (DP) as a powerful cryptographic primitive has achieved satisfying results in deep learning training, the existing schemes still fall short in preserving model utility, i.e., they either invoke a high noise scale or inevitably harm the original gradients. To address the above issues, in this paper, we present a more robust and provably secure approach for differentially private training called GReDP. Specifically, we compute the model gradients in the frequency domain and adopt a new approach to reduce the noise level. Unlike previous work, our GReDP only requires half of the noise scale compared to DPSGD [1] while keeping all the gradient information intact. We present a detailed analysis of our method both theoretically and empirically. The experimental results show that our GReDP works consistently better than the baselines on all models and training settings.

cs.CR↗