Searcharxiv⌕ Search

arXiv subjects

Yuan Lin

Publications and source records attributed to Yuan Lin.

At least 37 records · Page 2Linked to original sources

Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

We introduce Tarsier2, a state-of-the-art large vision-language model (LVLM) designed for generating detailed and accurate video descriptions, while also exhibiting superior general video understanding capabilities. Tarsier2 achieves significant advancements through three key upgrades: (1) Scaling pre-training data from 11M to 40M video-text pairs, enriching both volume and diversity; (2) Performing fine-grained temporal alignment during supervised fine-tuning; (3) Using model-based sampling to automatically construct preference data and applying DPO training for optimization. Extensive experiments show that Tarsier2-7B consistently outperforms leading proprietary models, including GPT-4o and Gemini 1.5 Pro, in detailed video description tasks. On the DREAM-1K benchmark, Tarsier2-7B improves F1 by 2.8% over GPT-4o and 5.8% over Gemini-1.5-Pro. In human side-by-side evaluations, Tarsier2-7B shows a +8.6% performance advantage over GPT-4o and +24.9% over Gemini-1.5-Pro. Tarsier2-7B also sets new state-of-the-art results across 15 public benchmarks, spanning tasks such as video question-answering, video grounding, hallucination test, and embodied question-answering, demonstrating its versatility as a robust generalist vision-language model.

cs.CV↗

AutoPatent: A Multi-Agent Framework for Automatic Patent Generation

As the capabilities of Large Language Models (LLMs) continue to advance, the field of patent processing has garnered increased attention within the natural language processing community. However, the majority of research has been concentrated on classification tasks, such as patent categorization and examination, or on short text generation tasks like patent summarization and patent quizzes. In this paper, we introduce a novel and practical task known as Draft2Patent, along with its corresponding D2P benchmark, which challenges LLMs to generate full-length patents averaging 17K tokens based on initial drafts. Patents present a significant challenge to LLMs due to their specialized nature, standardized terminology, and extensive length. We propose a multi-agent framework called AutoPatent which leverages the LLM-based planner agent, writer agents, and examiner agent with PGTree and RRAG to generate lengthy, intricate, and high-quality complete patent documents. The experimental results demonstrate that our AutoPatent framework significantly enhances the ability to generate comprehensive patents across various LLMs. Furthermore, we have discovered that patents generated solely with the AutoPatent framework based on the Qwen2.5-7B model outperform those produced by larger and more powerful LLMs, such as GPT-4o, Qwen2.5-72B, and LLAMA3.1-70B, in both objective metrics and human evaluations. We will make the data and code available upon acceptance at \url{https://github.com/QiYao-Wang/AutoPatent}.

cs.CL↗

AGILE: A Novel Reinforcement Learning Framework of LLM Agents

We introduce a novel reinforcement learning framework of LLM agents named AGILE (AGent that Interacts and Learns from Environments) designed to perform complex conversational tasks with users, leveraging LLMs, memory, tools, and interactions with experts. The agent possesses capabilities beyond conversation, including reflection, tool usage, and expert consultation. We formulate the construction of such an LLM agent as a reinforcement learning (RL) problem, in which the LLM serves as the policy model. We fine-tune the LLM using labeled data of actions and the PPO algorithm. We focus on question answering and release a dataset for agents called ProductQA, comprising challenging questions in online shopping. Our extensive experiments on ProductQA, MedMCQA and HotPotQA show that AGILE agents based on 7B and 13B LLMs trained with PPO can outperform GPT-4 agents. Our ablation study highlights the indispensability of memory, tools, consultation, reflection, and reinforcement learning in achieving the agent's strong performance. Datasets and code are available at https://github.com/bytarnish/AGILE.

cs.LG↗

Scalable Reshaping of Diamond Particles via Programmable Nanosculpting

Diamond particles have many interesting properties and possible applications. However, producing diamond particles with well-defined shapes at scale is challenging because diamonds are chemically inert and extremely hard. Here, we show air oxidation, a routine method for purifying diamonds, can be used to precisely shape diamond particles at scale. By exploiting the distinct reactivities of different crystal facets and defects inside the diamond, layer-by-layer outward-to-inward and inward-to-outward oxidation produced diverse diamond shapes including sphere, twisted surface, pyramidal islands, inverted pyramids, nano-flowers, and hollow polygons. The nanosculpted diamonds had more and finer features that enabled them to outperform the original raw diamonds in various applications. Using experimental observations and Monte Carlo simulations, we built a shape library that guides the design and fabrication of diamond particles with well-defined shapes and functional value. Our study presents a simple, economical and scalable way to produce shape-customized diamonds for various photonics, catalysis, quantum and information technology applications.

cond-mat.mtrl-sci↗

A diamond heater-thermometer microsensor for measuring localized thermal conductivity: a case study in gelatin hydrogel

Understanding the microscopic thermal effects of the hydrogel is important for its application in diverse fields, including thermal-related studies in tissue engineering and thermal management for flexible electronic devices. In recent decades, localized thermal properties, such as thermal conductivity, have often been overlooked due to technical limitations. To tackle this, we propose a new hybrid diamond microsensor that is capable of simultaneous temperature control and readout in a decoupled manner. Specifically, the sensor consists of a silicon pillar (heater) at about 10 microns in length, topped by a micron-sized diamond particle that contains silicon-vacancy (SiV) centers (thermometer) with 1.29 K*Hz^(-1/2) temperature measurement sensitivity. Combining this innovative, scalable sensor with a newly established simulation model that can transform heating-laser-induced temperature change into thermal conductivity, we introduced an all-optical decoupled method with about 0.05 W/(m* K) precision, which can reduce laser crosstalk. For the first time, we track the thermal conductivity change of hydrogels during the gelation process and demonstrate the existence of variation. We introduce a rapid, undisturbed technique for measuring microscale thermal conductivity, potentially serving as a valuable tool for cellular thermometry and highlight the idea that decoupling can reduce crosstalk from different lasers, which is helpful for quantum sensing.

physics.optics↗

Discriminative Addressing of Versatile Nanodiamonds via Physically-Enabled Classifier in Complex Bio-Systems

Nitrogen-vacancy (NV) centers show great potentials for nanoscale bio-sensing and bio-imaging. Nevertheless, their envisioned bio-applications suffer from intrinsic background noise due to unavoidable light scattering and autofluorescence in cells and tissues. Herein, we develop a novel all-optical modulated imaging method via physically-enabled classifier, for on-demand and direct access to NV fluorescence at pixel resolution while effectively filtering out background noise. Specifically, NV fluorescence can be modulated optically to exhibit sinusoid-like variations, providing basis for classification. We validate our method in various complex biological scenarios with fluorescence interference, ranging from cells to organisms. Notably, our classification-based approach achieves almost 10^6 times enhancement of signal-to-background ratio (SBR) for fluorescent nanodiamonds (FNDs) in neural protein imaging. We also demonstrate 4-fold contrast improvement in optically-detected magnetic resonance measurements (ODMR) of FNDs inside stained cells. Our technique offers a generic, explainable and robust solution, applicable for realistic high-fidelity imaging and sensing in challenging noise-laden scenarios.

physics.optics↗

AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering

We propose a novel and challenging benchmark, AutoEval-Video, to comprehensively evaluate large vision-language models in open-ended video question answering. The comprehensiveness of AutoEval-Video is demonstrated in two aspects: 1) AutoEval-Video constructs open-ended video-questions across 9 skill dimensions, addressing capabilities of perception, comprehension, and generation. 2) AutoEval-Video contains newly collected videos that cover over 40 distinct themes. To efficiently evaluate responses to the open-ended questions, we employ an LLM-based evaluation approach, but instead of merely providing a reference answer, we annotate unique evaluation rules for every single instance (video-question pair). To maximize the robustness of these rules, we develop a novel adversarial annotation mechanism. By using instance-specific rules as prompt, GPT-4, as an automatic evaluator, can achieve a stable evaluation accuracy of around 97.0%, comparable to the 94.9% - 97.5% accuracy of a human evaluator. Furthermore, we assess the performance of eight large vision-language models on AutoEval-Video. Among them, GPT-4V(ision) significantly outperforms other models, achieving an accuracy of 32.2%. However, there is still substantial room for improvement compared to human accuracy of 72.8%. By conducting an extensive case study, we uncover several drawbacks of GPT-4V, such as limited temporal and dynamic comprehension, and overly general responses. Code is available at https://github.com/Xiuyuan-Chen/AutoEval-Video.

cs.CV↗

IPEval: A Bilingual Intellectual Property Agency Consultation Evaluation Benchmark for Large Language Models

The rapid development of Large Language Models (LLMs) in vertical domains, including intellectual property (IP), lacks a specific evaluation benchmark for assessing their understanding, application, and reasoning abilities. To fill this gap, we introduce IPEval, the first evaluation benchmark tailored for IP agency and consulting tasks. IPEval comprises 2657 multiple-choice questions across four major dimensions: creation, application, protection, and management of IP. These questions span patent rights (inventions, utility models, designs), trademarks, copyrights, trade secrets, and other related laws. Evaluation methods include zero-shot, 5-few-shot, and Chain of Thought (CoT) for seven LLM types, predominantly in English or Chinese. Results show superior English performance by models like GPT series and Qwen series, while Chinese-centric LLMs excel in Chinese tests, albeit specialized IP LLMs lag behind general-purpose ones. Regional and temporal aspects of IP underscore the need for LLMs to grasp legal nuances and evolving laws. IPEval aims to accurately gauge LLM capabilities in IP and spur development of specialized models. Website: \url{https://ipeval.github.io/}

cs.CL↗

Mixed-Integer Optimal Control via Reinforcement Learning: A Case Study on Hybrid Electric Vehicle Energy Management

Many optimal control problems require the simultaneous output of discrete and continuous control variables. These problems are usually formulated as mixed-integer optimal control (MIOC) problems, which are challenging to solve due to the complexity of the solution space. Numerical methods such as branch-and-bound are computationally expensive and undesirable for real-time control. This paper proposes a novel hybrid-action reinforcement learning (HARL) algorithm, twin delayed deep deterministic actor-Q (TD3AQ), for MIOC problems. TD3AQ combines the advantages of both actor-critic and Q-learning methods, and can handle the discrete and continuous action spaces simultaneously. The proposed algorithm is evaluated on a plug-in hybrid electric vehicle (PHEV) energy management problem, where real-time control of the discrete variables, clutch engagement/disengagement and gear shift, and continuous variable, engine torque, is essential to maximize fuel economy while satisfying driving constraints. Simulation outcomes demonstrate that TD3AQ achieves control results close to optimality when compared with dynamic programming (DP), with just 4.69% difference. Furthermore, it surpasses the performance of baseline reinforcement learning algorithms.

eess.SY↗

Type-printable photodetector arrays for multichannel meta-infrared imaging

Multichannel meta-imaging, inspired by the parallel-processing capability of neuromorphic computing, offers significant advancements in resolution enhancement and edge discrimination in imaging systems, extending even into the mid- to far-infrared spectrum. Currently typical multichannel infrared imaging systems consist of separating optical gratings or merging multi-cameras, which require complex circuit design and heavy power consumption, hindering the implementation of advanced human-eye-like imagers. Here, we present a novel approach for printable graphene plasmonic photodetector arrays driven by a ferroelectric superdomain for multichannel meta-infrared imaging with enhanced edge discrimination. The fabricated photodetectors exhibited multiple spectral responses with zero-bias operation by directly rescaling the ferroelectric superdomain instead of reconstructing the separated gratings. We also demonstrated enhanced and faster shape classification (98.1%) and edge detection (98.2%) using our multichannel infrared images compared with single-channel detectors. Our proof-of-concept photodetector arrays simplify multichannel infrared imaging systems and hold great potential for applications in efficient edge detection in human-brain-type machine vision.

physics.optics↗

Autonomous vehicle decision and control through reinforcement learning with traffic flow randomization

Most of the current studies on autonomous vehicle decision-making and control tasks based on reinforcement learning are conducted in simulated environments. The training and testing of these studies are carried out under rule-based microscopic traffic flow, with little consideration of migrating them to real or near-real environments to test their performance. It may lead to a degradation in performance when the trained model is tested in more realistic traffic scenes. In this study, we propose a method to randomize the driving style and behavior of surrounding vehicles by randomizing certain parameters of the car-following model and the lane-changing model of rule-based microscopic traffic flow in SUMO. We trained policies with deep reinforcement learning algorithms under the domain randomized rule-based microscopic traffic flow in freeway and merging scenes, and then tested them separately in rule-based microscopic traffic flow and high-fidelity microscopic traffic flow. Results indicate that the policy trained under domain randomization traffic flow has significantly better success rate and calculative reward compared to the models trained under other microscopic traffic flows.

eess.SY↗

Highway Discretionary Lane-change Decision and Control Using Model Predictive Control

To enable autonomous vehicles to perform discretionary lane change amidst the random traffic flow on highways, this paper introduces a decision-making and control method for vehicle lane change based on Model Predictive Control (MPC). This approach divides the driving control of vehicles on highways into two parts: lane-change decision and lane-change control, both of which are solved using the MPC method. In the lanechange decision module, the minimum driving costs for each lane are computed and compared by solving the MPC problem to make lane-change decisions. In the lane-change control module, a dynamic bicycle model is incorporated, and a multi-objective cost function is designed to obtain the optimal control inputs for the lane-change process. Additionally, A long-short term memory (LSTM) model is used to predict the trajectories of surrounding vehicles for both the MPC decision and control modules. The proposed lane-change decision and control method is simulated and validated in a driving simulator under random highway traffic conditions.

eess.SY↗

Steady-State Error Compensation for Reinforcement Learning with Quadratic Rewards

The selection of a reward function in Reinforcement Learning (RL) has garnered significant attention because of its impact on system performance. Issues of significant steady-state errors often manifest when quadratic reward functions are employed. Although absolute-value-type reward functions alleviate this problem, they tend to induce substantial fluctuations in specific system states, leading to abrupt changes. In response to this challenge, this study proposes an approach that introduces an integral term. By integrating this integral term into quadratic-type reward functions, the RL algorithm is adeptly tuned, augmenting the system's consideration of reward history, and consequently alleviates concerns related to steady-state errors. Through experiments and performance evaluations on the Adaptive Cruise Control (ACC) and lane change models, we validate that the proposed method effectively diminishes steady-state errors and does not cause significant spikes in some system states.

eess.SY↗

Discretionary Lane-Change Decision and Control via Parameterized Soft Actor-Critic for Hybrid Action Space

This study focuses on a crucial task in the field of autonomous driving, autonomous lane change. Autonomous lane change plays a pivotal role in improving traffic flow, alleviating driver burden, and reducing the risk of traffic accidents. However, due to the complexity and uncertainty of lane-change scenarios, the functionality of autonomous lane change still faces challenges. In this research, we conducted autonomous lane-change simulations using both deep reinforcement learning (DRL) and model predictive control (MPC). Specifically, we used the parameterized soft actor--critic (PASAC) algorithm to train a DRL-based lane-change strategy to output both discrete lane-change decisions and continuous longitudinal vehicle acceleration. We also used MPC for lane selection based on the smallest predictive car-following costs for the different lanes. For the first time, we compared the performance of DRL and MPC in the context of lane-change decisions. The simulation results indicated that, under the same reward/cost function and traffic flow, both MPC and PASAC achieved a collision rate of 0%. PASAC demonstrated a comparable performance to MPC in terms of average rewards/costs and vehicle speeds.

cs.RO↗

Plug-in Hybrid Electric Vehicle Energy Management with Clutch Engagement Control via Continuous-Discrete Reinforcement Learning

Energy management strategy (EMS) is a key technology for plug-in hybrid electric vehicles (PHEVs). The energy management of certain series-parallel PHEVs involves the control of continuous variables, such as engine torque, and discrete variables, such as clutch engagement/disengagement. We establish a control-oriented model for a series-parallel plug-in hybrid system with clutch engagement control from the perspective of mixed-integer programming. Subsequently, we design an EMS based on continuous-discrete reinforcement learning (CDRL), which enables simultaneous output of continuous and discrete variables. During training, we introduce state-of-charge (SOC) randomization to ensure that the hybrid system exhibits optimal energy-saving performance in both high and low SOC. Finally, the effectiveness of the proposed CDRL strategy is verified by comparing EMS based on charge-depleting charge-sustaining (CD-CS) with rule-based clutch engagement control, and Dynamic Programming (DP). The simulation results show that, under a high SOC, the CDRL strategy proposed in this paper can improve energy efficiency by 8.3% compared to CD-CS, and the energy consumption is just 6.6% higher than the global optimum based on DP, while under a low SOC, the numbers are 4.1% and 3.9%, respectively.

eess.SY↗

Safe Hybrid-Action Reinforcement Learning-Based Decision and Control for Discretionary Lane Change

Autonomous lane-change, a key feature of advanced driver-assistance systems, can enhance traffic efficiency and reduce the incidence of accidents. However, safe driving of autonomous vehicles remains challenging in complex environments. How to perform safe and appropriate lane change is a popular topic of research in the field of autonomous driving. Currently, few papers consider the safety of reinforcement learning in autonomous lane-change scenarios. We introduce safe hybrid-action reinforcement learning into discretionary lane change for the first time and propose Parameterized Soft Actor-Critic with PID Lagrangian (PASAC-PIDLag) algorithm. Furthermore, we conduct a comparative analysis of the Parameterized Soft Actor-Critic (PASAC), which is an unsafe version of PASAC-PIDLag. Both algorithms are employed to train the lane-change strategy of autonomous vehicles to output discrete lane-change decision and longitudinal vehicle acceleration. Our simulation results indicate that at a traffic density of 15 vehicles per kilometer (15 veh/km), the PASAC-PIDLag algorithm exhibits superior safety with a collision rate of 0%, outperforming the PASAC algorithm, which has a collision rate of 1%. The outcomes of the generalization assessments reveal that at low traffic density levels, both the PASAC-PIDLag and PASAC algorithms are proficient in attaining a 0% collision rate. Under conditions of high traffic flow density, the PASAC-PIDLag algorithm surpasses PASAC in terms of both safety and optimality.

cs.RO↗

Bi-Preference Learning Heterogeneous Hypergraph Networks for Session-based Recommendation

Session-based recommendation intends to predict next purchased items based on anonymous behavior sequences. Numerous economic studies have revealed that item price is a key factor influencing user purchase decisions. Unfortunately, existing methods for session-based recommendation only aim at capturing user interest preference, while ignoring user price preference. Actually, there are primarily two challenges preventing us from accessing price preference. Firstly, the price preference is highly associated to various item features (i.e., category and brand), which asks us to mine price preference from heterogeneous information. Secondly, price preference and interest preference are interdependent and collectively determine user choice, necessitating that we jointly consider both price and interest preference for intent modeling. To handle above challenges, we propose a novel approach Bi-Preference Learning Heterogeneous Hypergraph Networks (BiPNet) for session-based recommendation. Specifically, the customized heterogeneous hypergraph networks with a triple-level convolution are devised to capture user price and interest preference from heterogeneous features of items. Besides, we develop a Bi-Preference Learning schema to explore mutual relations between price and interest preference and collectively learn these two preferences under the multi-task learning architecture. Extensive experiments on multiple public datasets confirm the superiority of BiPNet over competitive baselines. Additional research also supports the notion that the price is crucial for the task.

cs.IR↗

Dynamic Modulation of Electromagnetically Induced Transparency Metamaterials through Mode Coupling and Stretchable Design

The active control of electromagnetically induced transparency (EIT) metamaterials (MM) has the potential to revolutionize communication networks without relying on quantum technology. However, current reconfigurable systems offer limited flexibility and have high fabrication costs and difficulties. In this study, we examine a classical EIT metamaterial and discover a novel modulation mechanism that leverages mode coupling to dynamically adjust the bandwidth and group delay of the EIT MM. This mechanism is verified through analyses of the electric field and surface charge density distributions. Additionally, a robust coupled Lorentz oscillator model is used to explain the coupling mechanism, with results that are in good agreement with simulations and experiments. To capitalize on this mechanism, we propose a block-definition approach where the MM is divided into stretchable sections, allowing for dynamic modulation of the bandwidth and group delay by stretching the EIT MM. Furthermore, the fabrication process is highly compatible with traditional flexible printed circuit board techniques. Our block-definition EIT MM offers unprecedented tunability and flexibility, requiring no complex components or specialized materials, making it a promising candidate for tunable slow-wave devices and other reconfigurable microwave applications.

physics.app-ph↗