SearcharxivSearch

arXiv subjects

Guangrui Li

Publications and source records attributed to Guangrui Li.

7 recordsLinked to original sources

PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants

Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This changes the unit of evaluation: a fluent response can still fail by ignoring shopper context, contradicting itself across components, or leaving defects too vague to localize. Existing personalization, grounding, and LLM-as-a-judge benchmarks cover pieces of this problem, but they do not define a joint evaluation target for structured shopping-assistant responses. We formulate this missing evaluation target as PACE: Personalized, Actionable, Compositional, and Evidence-grounded evaluation. We instantiate PACE with two artifacts: PACEShop, a benchmark dataset that makes the target measurable through 22,625 controlled records with structured personas, auditable evidence pools, GOOD/BAD labels, and gold defect family and location annotations; and PACEJudge, a training-free judging protocol that makes the target reportable through a structured output contract. Our experiments show that generic judges can recognize broad quality but fail to recover the diagnostic fields required for PACE; PACEShop makes these failures verifiable, and PACEJudge improves persona-source, cross-component, grounding, and family/location closure without retraining, showing that realistic shopping-assistant evaluation requires a task-matched output contract rather than only a stronger backbone or scalar prompt.

cs.CL

The World Won't Stay Still: Programmable Evolution for Agent Benchmarks

LLM-powered tool-calling agents fulfill user requests by interacting with environments, querying data, and invoking tools in a multi-turn process. Yet, most existing benchmarks evaluate these systems under static environment interfaces, with fixed schemas and toolsets, making it difficult to assess how agents behave as environments evolves -- when capabilities are added, reorganized, or deprecated across successive environment versions. In this paper, we study structured environment evolution as a benchmark-construction problem for tool-calling agents. We propose ProEvolve, a graph-based framework that makes environment evolution programmable. At its core, a typed relational graph provides a unified, explicit representation of the environment - data, tools, and schema. Under this formalism, adding, removing, or modifying capabilities are expressed as graph transformations that coherently propagate updates across tools, schemas, and data access. Building on this, ProEvolve supports (1) automatic generation of evolved executable environments through explicit graph transformations, and (2) graph-grounded construction of task sandboxes via subgraph sampling and instantiation. We validate ProEvolve in two tool-calling domains, e-commerce and airline booking, in terms of quality, implementation validity, and failure modes. Finally, we use the generated benchmark as a downstream diagnostic to study how representative agents behave under structured environment evolution.

cs.AI

Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding

Applying Multimodal Large Language Models (MLLMs) to video understanding presents significant challenges due to the need to model temporal relations across frames. Existing approaches adopt either implicit temporal modeling, relying solely on the LLM decoder, or explicit temporal modeling, employing auxiliary temporal encoders. To investigate this debate between the two paradigms, we propose the Stackable Temporal Encoder (STE). STE enables flexible explicit temporal modeling with adjustable temporal receptive fields and token compression ratios. Using STE, we systematically compare implicit and explicit temporal modeling across dimensions such as overall performance, token compression effectiveness, and temporal-specific understanding. We also explore STE's design considerations and broader impacts as a plug-in module and in image modalities. Our findings emphasize the critical role of explicit temporal modeling, providing actionable insights to advance video MLLMs.

cs.CV

Coherent Power Scaling in Photonic Crystal Surface Emitting Laser Arrays

A key benefit of photonic crystal surface emitting lasers (PCSELs) is the abillity to increase output power through scaling the emission area while mainting high quality single mode emission, allowing them to close the brightness gap which exists between semiconductor lasers and gas and fibre lasers. However, there are practical limits to the size, and hence power, of an individual PCSEL device and there are trade-offs between single-mode stability and parasitic in-plane losses with increasing device size. In this paper we discuss 2D coherent arrays as an approach to area and coherent power scaling of PCSELs. We demonstrate in two and three element PCSEL arrays an increase in the differential efficiency of the system due to a reduction in in-plane loss.

physics.optics

An improved spectrophotometry tests the Einstein-Smoluchowski equation: a revisit and update

Light-matter interaction in solvents has attracted continuous attention for theoretical prediction and experimental measure, due in part to the simple curiosity to nature, and in part to increasing calls from solvent-involved applications. Yet hitherto, a majority of reliable spectrophotometric measurements on transparent solvents upon visible light end up using long-path-length cells, usually over dozens of cm, rendering the measures costly and complex; meanwhile, the guidance for choosing the best formula to describe solvent scattering has remained unsettled. Here we theoretically and experimentally demonstrate a simple, low-cost, and versatile spectrophotometric method, recording sensitivity 10-4 dB/cm over 0.5 cm differential path length based on using a standard double-beam spectrophotometer. We attest the method reduces the path length by a factor of 100 while still making its closest approach to the record-low measurements. Revisiting the present equations of solvent scattering, we unfold that they all give similar-predictive-values, revealing the criterion of choice merely on the formula's simple practicality. Following the clarification of wavelengths over which light scattering dictates the solvent's extinction, we identify that the discrepancies persist between the calculated scattering coefficients and those measured results, suggesting the need for improving solvent scattering theory to comprehend the phenomenon in greater depth.

physics.chem-ph

Concept for a Future Super Proton-Proton Collider

Following the discovery of the Higgs boson at LHC, new large colliders are being studied by the international high-energy community to explore Higgs physics in detail and new physics beyond the Standard Model. In China, a two-stage circular collider project CEPC-SPPC is proposed, with the first stage CEPC (Circular Electron Positron Collier, a so-called Higgs factory) focused on Higgs physics, and the second stage SPPC (Super Proton-Proton Collider) focused on new physics beyond the Standard Model. This paper discusses this second stage.

physics.acc-ph