SearcharxivSearch

arXiv subjects

Rui Shen

Publications and source records attributed to Rui Shen.

16 recordsLinked to original sources

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fragmented architectures, cascaded pipelines, and misaligned representation spaces. We argue that this divide is not merely an engineering artifact, but a structural limitation that hinders the emergence of native multimodal intelligence. Hence, we introduce SenseNova-U1, a native unified multimodal paradigm built upon NEO-unify, in which understanding and generation evolve as synergistic views of a single underlying process. We launch two native unified variants, SenseNova-U1-8B-MoT and SenseNova-U1-A3B-MoT, built on dense (8B) and mixture-of-experts (30B-A3B) understanding baselines, respectively. Designed from first principles, they rival top-tier understanding-only VLMs across text understanding, vision-language perception, knowledge reasoning, agentic decision-making, and spatial intelligence. Meanwhile, they deliver strong semantic consistency and visual fidelity, excelling in conventional or knowledge-intensive any-to-image (X2I) synthesis, complex text-rich infographic generation, and interleaved vision-language generation, with or without think patterns. Beyond performance, we show detailed model design, data preprocessing, pre-/post-training, and inference strategies to support community research. Last but not least, preliminary evidence demonstrates that our models extend beyond perception and generation, performing strongly in vision-language-action (VLA) and world model (WM) scenarios. This points toward a broader roadmap where models do not translate between modalities, but think and act across them in a native manner. Multimodal AI is no longer about connecting separate systems, but about building a unified one and trusting the necessary capabilities to emerge from within.

cs.CV

ACPO: Counteracting Likelihood Displacement in Vision-Language Alignment with Asymmetric Constraints

While Direct Preference Optimization (DPO) has become the de facto approach for aligning Large Vision-Language Models (LVLMs), it suffers from Likelihood Displacement, where the probability of both chosen and rejected responses collapses. This optimization flaw is especially detrimental in multimodal settings: the erosion of chosen likelihoods -- a failure we term Visual Anchor Collapse -- causes models to abandon visual evidence for strong language priors, precipitating significant hallucinations. To address this, we propose Asymmetric Constrained Preference Optimization (ACPO), a modality-agnostic alignment mechanism that applies dynamic, target-oriented scaling to preference optimization. ACPO derives a complexity-aware scaling coefficient applied exclusively to the rejected reward, asymmetrically suppressing the gradient flow on the rejected term while preserving the chosen distribution as a gradient-stable reference. While fundamentally a general-purpose objective, breaking this gradient symmetry is crucial for multimodal tasks, as it mitigates the suppression of visual tokens by language priors. Experiments on InternVL models demonstrate that ACPO effectively reverses the chosen-reward degradation of standard DPO. By halting Visual Anchor Collapse, ACPO generally outperforms baselines on hallucination benchmarks (HallusionBench, MM-IFEval) and general leaderboards (MMBench, MMStar, OCRBenchV2) while driving concurrent improvements in general capabilities.

cs.CV

Monolithic integration of diverse crystalline thin films on diamond for near-junction thermal management

The pursuit of extreme miniaturization and high power in 6G RF front-ends has cast thermal dissipation as the central challenge. Here, we have demonstrated the monolithic integration of functionally distinct single-crystal thin films, including \b{eta}-Ga2O3, Si, GaN, and LiTaO3, onto a single diamond substrate using a multi-step transfer printing technique. Focusing on the critical \b{eta}-Ga2O3/diamond interface, we achieve an exceptional interfacial thermal conductance (ITC) of 149 MW m-2 K-1 through ultra-high vacuum (UHV) annealing, creating an atomically sharp interface featuring covalent bonding. Vibrational electron energy-loss spectroscopy (EELS) analysis combining with molecular dynamics (MD) simulations reveal that distinctive interfacial phonon modes at the \b{eta}-Ga2O3/diamond heterointerface dominate ultrahigh ITC. We experimentally demonstrate that by improving the ITC, the thermal resistance (Rth) of a diamond-based \b{eta}-Ga2O3 MOSFET is driven to a record-low value of 1.58 K mm W-1, underscoring the critical role of interface engineering in near-junction thermal management for diamond-integrated devices. This work demonstrates a scalable, diamond-based monolithic integration platform designed to solve the near-junction thermal challenges in high-power RF front-ends.

cond-mat.mtrl-sci

Electrically pumped AlGaN edge-emitting UV-B laser diodes grown by molecular beam epitaxy

Mid and deep ultraviolet (UV) laser diodes remain among the least explored devices in semiconductor optoelectronics, despite their importance for spectroscopy, biochemical sensing, disinfection, and emerging quantum photonics. Here, we demonstrate an electrically pumped AlGaN-based laser diode operating in the UV-B band (280-315 nm). The device is grown by molecular beam epitaxy (MBE) on single-crystal AlN substrate and fabricated in a ridge-waveguide geometry. The laser diode operates at 298.5 nm and exhibits a relatively low threshold current density of 3.4 kA/cm$^2$. Clear nonlinear light-current characteristics and pronounced spectral narrowing with a full-width-at-half-maximum (FWHM) of 0.2 nm are measured above threshold.

physics.optics

Improving Public Service Chatbot Design and Civic Impact: Investigation of Citizens' Perceptions of a Metro City 311 Chatbot

As governments increasingly adopt digital tools, public service chatbots have emerged as a growing communication channel. This paper explores the design considerations and engagement opportunities of public service chatbots, using a 311 chatbot from a metropolitan city as a case study. Our qualitative study consisted of official survey data and 16 interviews examining stakeholder experiences and design preferences for the chatbot. We found two key areas of concern regarding these public chatbots: individual-level and community-level. At the individual level, citizens experience three key challenges: interpretation, transparency, and social contextualization. Moreover, the current chatbot design prioritizes the efficient completion of individual tasks but neglects the broader community perspective. It overlooks how individuals interact and discuss problems collectively within their communities. To address these concerns, we offer design opportunities for creating more intelligent, transparent, community-oriented chatbots that better engage individuals and their communities.

cs.HC

Enhanced oil recovery in reservoirs via diffusion-driven $\text{CO}_{2}$ flooding: Experimental insights and material balance modeling

$\text{CO}_{2}$ flooding is central to carbon utilization technologies, yet conventional waterflooding models fail to capture the complex interactions between CO$_2$ and formation fluids. In this study, one- and two-dimensional nuclear magnetic resonance experiments reveal that $\text{CO}_{2}$ markedly enhances crude oil mobility during miscible displacement via multiple synergistic mechanisms, yielding a recovery factor of $60.97\%$, which surpasses that of immiscible displacement (maximum $57.53\%$). Guided by these findings, we propose a convection-diffusion model that incorporates the diffusion coefficient ($D$) and porosity ($\phi$) as key parameters. This model captures the spatiotemporal evolution of the $\text{CO}_{2}$ front and addresses a key limitation of conventional formulations-the omission of diffusion effects. It improves predictions of gas breakthrough time and enables optimized injection design for low-permeability reservoirs. Extending classical material balance theory, we develop an enhanced $\text{CO}_{2}$ flooding equation that integrates critical transport phenomena. This formulation incorporates $\text{CO}_{2}$ diffusion, oil phase expansion, reservoir adsorption, and gas compressibility to describe the dynamic transport and mass compensation of injected $\text{CO}_{2}$. Validation through experimental and numerical data confirms the model's robustness and applicability under low-permeability conditions. The proposed framework overcomes limitations of physical experiments under extreme environments and offers theoretical insight into oil recovery enhancement and $\text{CO}_{2}$ injection strategy optimization.

physics.flu-dyn

P-FOLIO: Evaluating and Improving Logical Reasoning with Abundant Human-Written Reasoning Chains

Existing methods on understanding the capabilities of LLMs in logical reasoning rely on binary entailment classification or synthetically derived rationales, which are not sufficient for proper investigation of model's capabilities. We present P-FOLIO, a human-annotated dataset consisting of diverse and complex reasoning chains for a set of realistic logical reasoning stories also written by humans. P-FOLIO is collected with an annotation protocol that facilitates humans to annotate well-structured natural language proofs for first-order logic reasoning problems in a step-by-step manner. The number of reasoning steps in P-FOLIO span from 0 to 20. We further use P-FOLIO to evaluate and improve large-language-model (LLM) reasoning capabilities. We evaluate LLM reasoning capabilities at a fine granularity via single-step inference rule classification, with more diverse inference rules of more diverse and higher levels of complexities than previous works. Given that a single model-generated reasoning chain could take a completely different path than the human-annotated one, we sample multiple reasoning chains from a model and use pass@k metrics for evaluating the quality of model-generated reasoning chains. We show that human-written reasoning chains significantly boost the logical reasoning capabilities of LLMs via many-shot prompting and fine-tuning. Furthermore, fine-tuning Llama3-7B on P-FOLIO improves the model performance by 10% or more on three other out-of-domain logical reasoning datasets. We also conduct detailed analysis to show where most powerful LLMs fall short in reasoning. We will release the dataset and code publicly.

cs.AI

Blockchain-based Federated Recommendation with Incentive Mechanism

Nowadays, federated recommendation technology is rapidly evolving to help multiple organisations share data and train models while meeting user privacy, data security and government regulatory requirements. However, federated recommendation increases customer system costs such as power, computational and communication resources. Besides, federated recommendation systems are also susceptible to model attacks and data poisoning by participating malicious clients. Therefore, most customers are unwilling to participate in federated recommendation without any incentive. To address these problems, we propose a blockchain-based federated recommendation system with incentive mechanism to promote more trustworthy, secure, and efficient federated recommendation service. First, we construct a federated recommendation system based on NeuMF and FedAvg. Then we introduce a reverse auction mechanism to select optimal clients that can maximize the social surplus. Finally, we employ blockchain for on-chain evidence storage of models to ensure the safety of the federated recommendation system. The experimental results show that our proposed incentive mechanism can attract clients with superior training data to engage in the federal recommendation at a lower cost, which can increase the economic benefit of federal recommendation by 54.9\% while improve the recommendation performance. Thus our work provides theoretical and technological support for the construction of a harmonious and healthy ecological environment for the application of federal recommendation.

cs.IR

Wonderful Team: Zero-Shot Physical Task Planning with Visual LLMs

We introduce Wonderful Team, a multi-agent Vision Large Language Model (VLLM) framework for executing high-level robotic planning in a zero-shot regime. In our context, zero-shot high-level planning means that for a novel environment, we provide a VLLM with an image of the robot's surroundings and a task description, and the VLLM outputs the sequence of actions necessary for the robot to complete the task. Unlike previous methods for high-level visual planning for robotic manipulation, our method uses VLLMs for the entire planning process, enabling a more tightly integrated loop between perception, control, and planning. As a result, Wonderful Team's performance on real-world semantic and physical planning tasks often exceeds methods that rely on separate vision systems. For example, we see an average 40% success rate improvement on VimaBench over prior methods such as NLaP, an average 30% improvement over Trajectory Generators on tasks from the Trajectory Generator paper, including drawing and wiping a plate, and an average 70% improvement over Trajectory Generators on a new set of semantic reasoning tasks including environment rearrangement with implicit linguistic constraints. We hope these results highlight the rapid improvements of VLLMs in the past year, and motivate the community to consider VLLMs as an option for some high-level robotic planning problems in the future.

cs.AI

Cold Diffusion on the Replay Buffer: Learning to Plan from Known Good States

Learning from demonstrations (LfD) has successfully trained robots to exhibit remarkable generalization capabilities. However, many powerful imitation techniques do not prioritize the feasibility of the robot behaviors they generate. In this work, we explore the feasibility of plans produced by LfD. As in prior work, we employ a temporal diffusion model with fixed start and goal states to facilitate imitation through in-painting. Unlike previous studies, we apply cold diffusion to ensure the optimization process is directed through the agent's replay buffer of previously visited states. This routing approach increases the likelihood that the final trajectories will predominantly occupy the feasible region of the robot's state space. We test this method in simulated robotic environments with obstacles and observe a significant improvement in the agent's ability to avoid these obstacles during planning.

cs.RO

L2CEval: Evaluating Language-to-Code Generation Capabilities of Large Language Models

Recently, large language models (LLMs), especially those that are pretrained on code, have demonstrated strong capabilities in generating programs from natural language inputs in a few-shot or even zero-shot manner. Despite promising results, there is a notable lack of a comprehensive evaluation of these models language-to-code generation capabilities. Existing studies often focus on specific tasks, model architectures, or learning paradigms, leading to a fragmented understanding of the overall landscape. In this work, we present L2CEval, a systematic evaluation of the language-to-code generation capabilities of LLMs on 7 tasks across the domain spectrum of semantic parsing, math reasoning and Python programming, analyzing the factors that potentially affect their performance, such as model size, pretraining data, instruction tuning, and different prompting methods. In addition to assessing model performance, we measure confidence calibration for the models and conduct human evaluations of the output programs. This enables us to identify and analyze the typical failure modes across various tasks and models. L2CEval offers a comprehensive understanding of the capabilities and limitations of LLMs in language-to-code generation. We also release the evaluation framework and all model outputs, hoping to lay the groundwork for further future research in this domain.

cs.CL

Well-separated soliton-antisoliton pairs with an adjoint Higgs field in 4D space

We present single soliton states and soliton-antisoliton states with an adjoint Higgs field in 4D flat space. The action of a single soliton state diverges, while the action of soliton-antisoliton states converges. This means such solitons can exist in soliton-antisoliton states, although they cannot exist individually. The interaction in a soliton-antisoliton state takes a logarithmic dependence on separation. Such soliton-antisoliton states exhibit stability under a scaling transformation.

hep-th

Exact Penalty Algorithm of Strong Convertible Nonconvex Optimization

This paper defines a strong convertible nonconvex(SCN) function for solving the unconstrained optimization problems with the nonconvex or nonsmooth(nondifferentiable) function. First, many examples of SCN function are given, where the SCN functions are nonconvex or nonsmooth. Second, the operational properties of the SCN functions are proved, including addition, multiplication, compound operations and so on. Third, the SCN forms of some special functions common in machine learning and engineering applications are presented respectively where these SCN function optimization problems can be transformed into minmax problems with a convex and concave objective function. Fourth,a minmax optimization problem of SCN function and its penalty function are defined. The optimization condition,exactness and stability of the minmax optimization problem are proved. Finally, an algorithm of penalty function to solve the minmax optimization problem and its convergence are given. This paper provides an efficient technique for solving unconstrained nonconvex or nonsmooth(nondifferentiable) optimization problems to avoid using subdifferentiation or smoothing techniques.

math.OC

Optimization Condition and Algorithm of Optimization with Convertible Nonconvex Function

The paper introduces several new concepts for solving nonconvex or nonsmooth optimization problems, including convertible nonconvex function, exact convertible nonconvex function and differentiable convertible nonconvex function. It is proved herein many nonconvex functions or nonsmooth (or discontinuous) functions are actually convertible nonconvex functions and convertible nonconvex function operations such as addition, subtraction, multiplication or division result in convertible nonconvex functions. The sufficient condition for judging a global optimal solution to unconstrained optimization problems with differentiable convertible nonconvex functions is proved, which is equivalent to Karush-Kuhn-Tucker(KKT) condition. Two Lagrange functions of differentiable convertible nonconvex function are defined with their dual problems defined accordingly. The strong duality theorem is proved, showing that the optimal objective value of the global optimal solution is equal to the optimal objective value of the dual problem, which is equivalent to KKT condition. An augmented Lagrangian penalty function algorithm is proposed and its convergence is proved. So the paper provides a new idea for solving unconstrained nonconvex or non-smooth optimization problems and avoids subdifferentiation or smoothing techniques by using some gradient search algorithms, such as gradient descent algorithm, Newton algorithm and so on.

math.OC

Traffic flow brake light model simulation based on driver behavior learning

The theory of urban traffic flow has been developed and new types of meta-automata have emerged and simulate realistic traffic conditions relatively well. Among these models, the brake light model can simulate the three-phase traffic flow theory very well. However, the existing brake light model also has certain shortcomings, in that the model will change the speed of congestion propagation upward when the model is covariant, which is not realistic, and the model also lacks simulation parameters for driver behavior. In this paper, we propose a new model based on the brake light model, which can achieve a certain degree of simulation of driver behavior by adjusting the parameters, and also achieve a stabilization of the propagation speed of the congestion wave when the parameters are changed.

math.CV

Floquet Weyl Semimetal Induced by Off-Resonant Light

We propose that a Floquet Weyl semimetal state can be induced in three-dimensional topological insulators, either nonmagnetic or magnetic, by the application of off-resonant light. The virtual photon processes play a critical role in renormalizing the Dirac mass and so resulting in a topological semimetal with vanishing gap at Weyl points. The present mechanism via off-resonant light is quite different from that via on-resonant light, the latter being recently suggested to give rise to a Floquet topological state in ordinary band insulators.

cond-mat.mes-hall