SearcharxivSearch

arXiv subjects

Haitao Leng

Publications and source records attributed to Haitao Leng.

13 recordsLinked to original sources

An Interface Green's Function Framework for Complete Discrete $W^{1,\infty}$ Analysis of Discontinuous Galerkin Methods

Pointwise error analysis of discontinuous Galerkin (DG) methods for the Poisson equation has received considerable attention during the past two decades. However, on convex polyhedral domains, existing analyses can only establish optimal error estimates in the broken $W^{1,\infty}$ seminorm for several DG methods. Since the broken $W^{1,\infty}$ seminorm does not control discontinuities across mesh interfaces, a complete discrete $W^{1,\infty}$ theory for discontinuous approximations on convex polyhedral domains has remained unavailable. In this paper, we develop an interface Green's function framework for the complete discrete $W^{1,\infty}$ analysis of DG methods on convex polyhedral domains. The proposed framework introduces new interface Green's functions that represent jumps of discontinuous approximations across mesh interfaces. Its central analytical ingredient is a new local energy estimate for these Green's functions, obtained by exploiting a cancellation between neighboring discrete delta functions. This estimate differs fundamentally from existing Green's function estimates and enables us to derive maximum-norm estimate for interface jumps without introducing additional logarithmic factors.

math.NA

MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation

Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medical vision-language models often lack precise localization, whereas medical segmenters typically rely on explicit target categories or precise spatial prompts. This divide is reinforced by a supervision mismatch: segmentation datasets provide precise masks but little language supervision, whereas medical vision-language data rarely pair language with dense spatial annotations. To address this gap, we present MedPixel, a unified medical pixel-language model built around a shared language--mask interface. To provide scalable supervision, we introduce MedPLG-440K, comprising approximately 440K pixel-language task samples constructed through a clinically motivated synthesis process without external LLM annotation. MedPixel is trained with joint multi-task supervised fine-tuning followed by Pixel-Level Preference Optimization, which uses ground-truth masks as offline verifiers to derive response preferences from mask quality. MedPixel supports a broad spectrum of tasks spanning explicit grounding, implicit reasoning, spatial interaction, grounded explanation, and medical VQA. Across this task spectrum, MedPixel achieves strong performance in both pixel-level prediction and response generation, together with effective zero-shot transfer to external grounding benchmarks and robustness to imperfect spatial prompts. Code and model checkpoints will be released at https://github.com/yhy-whu/Medpixel.

cs.CV

A unified analysis of maximum-norm estimates for a class of HDG methods for parabolic problem in polyhedral domains

This paper studies a general class of semi-discrete hybridizable discontinuous Galerkin (HDG) methods, including mixed methods, for parabolic problems in nonconvex polygonal and polyhedral domains. By developing local energy error estimates together with energy estimates for a regularized Green's function, we establish a unified framework to prove the maximum-norm stability of both the semigroup defined by the semi-discrete scheme and the corresponding discrete solutions. The stability analysis shows that the main challenges stem from the treatment of numerical flux variables and the inherent asymmetry of the discrete scheme. These challenges are intrinsic to numerical approaches formulated within the mixed framework for parabolic equations. The asymmetry prevents the direct application of the double kick-back argument, while the presence of flux variables requires special techniques to control their values at the initial time. We emphasize that the analytical tools developed here to address these challenges are sufficiently general to be adapted for maximum-norm stability investigations of a broader class of numerical schemes arising from mixed formulations of parabolic equations. Furthermore, by the stability results and their proofs, we derive the maximal regularity of the semi-discrete solution in $L^{\infty}((0,T);L^p(Ω))$-norm and reduce the maximum-norm error estimates to those of the corresponding elliptic equations and the $L^2$-orthogonal projection. Since the derivation of the local energy error estimates does not rely on the lifting operator, which is only available for simplicial meshes, our results (excluding mixed methods) remain valid for polygonal/polyhedral meshes.

math.NA

The Existence, uniqueness, and regularity of weak solutions for a thermodynamically consistent two-phase flow model in porous media

Thermodynamically consistent models for two-phase flow in porous media have attracted significant attention in recent years. In this paper, we prove the existence, uniqueness and regularity of the weak solution to such a recent model proposed in [25,35]. To this end, firstly, we introduce a fully implicit time semi-discrete approximation and a fully discrete approximation for an appropriate weak formulation of the thermodynamically consistent model. Next, by using the zeros of a vector field theorem, we prove the existence of the weak solution for the fully discrete approximation. Then the existence of weak solutions for the fully implicit time semi-discrete approximation and the weak formulation of the model are derived by the weak convergence technique and the energy stability estimate. Subsequently, by the Gr{\" o}nwall inequality, we prove the uniqueness result under the smoothness assumption on the chemical potential. Finally, combined with the regularity theory of elliptic partial differential equations (PDE), the regularity of the weak solution for the model with complete Neumann boundary conditions is established.

math.AP

Mix-Ecom: Towards Mixed-Type E-Commerce Dialogues with Complex Domain Rules

E-commerce agents contribute greatly to helping users complete their e-commerce needs. To promote further research and application of e-commerce agents, benchmarking frameworks are introduced for evaluating LLM agents in the e-commerce domain. Despite the progress, current benchmarks lack evaluating agents' capability to handle mixed-type e-commerce dialogue and complex domain rules. To address the issue, this work first introduces a novel corpus, termed Mix-ECom, which is constructed based on real-world customer-service dialogues with post-processing to remove user privacy and add CoT process. Specifically, Mix-ECom contains 4,799 samples with multiply dialogue types in each e-commerce dialogue, covering four dialogue types (QA, recommendation, task-oriented dialogue, and chit-chat), three e-commerce task types (pre-sales, logistics, after-sales), and 82 e-commerce rules. Furthermore, this work build baselines on Mix-Ecom and propose a dynamic framework to further improve the performance. Results show that current e-commerce agents lack sufficient capabilities to handle e-commerce dialogues, due to the hallucination cased by complex domain rules. The dataset will be publicly available.

cs.AI

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of multi-turn interaction, especially for long contexts. To address the issue, we first introduce a context modeling module, termed ContextQFormer, which utilizes a memory block to enhance the presentation of contextual information. Furthermore, to facilitate further research, we carefully build a new multi-turn multi-modal dialogue dataset (TMDialog) for pre-training, instruction-tuning, and evaluation, which will be open-sourced lately. Compared with other multi-modal dialogue datasets, TMDialog contains longer conversations, which supports the research of multi-turn multi-modal dialogue. In addition, ContextQFormer is compared with three baselines on TMDialog and experimental results illustrate that ContextQFormer achieves an improvement of 2%-4% in available rate over baselines.

cs.CL

GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art

Video Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural and contextual subtleties. Although Multimodal Large Language Models (MLLMs) and Chain-of-Thought (CoT) have demonstrated strong reasoning abilities in STEM tasks (e.g. mathematics and coding), they still struggle to generate creative expressions such as resonant jokes and insightful satire. Moreover, existing benchmarks are constrained by their limited modalities and insufficient categories, hindering the exploration of comprehensive creativity in video-based Comment Art creation. To address these limitations, we introduce GODBench, a novel benchmark that integrates video and text modalities to systematically evaluate MLLMs' abilities to compose Comment Art. Furthermore, inspired by the propagation patterns of waves in physics, we propose Ripple of Thought (RoT), a multi-step reasoning framework designed to enhance the creativity of MLLMs. Extensive experiments reveal that existing MLLMs and CoT methods still face significant challenges in understanding and generating creative video comments. In contrast, RoT provides an effective approach to improve creative composing, highlighting its potential to drive meaningful advancements in MLLM-based creativity. GODBench is publicly available at https://github.com/stan-lei/GODBench-ACL2025.

cs.CL

KwaiChat: A Large-Scale Video-Driven Multilingual Mixed-Type Dialogue Corpus

Video-based dialogue systems, such as education assistants, have compelling application value, thereby garnering growing interest. However, the current video-based dialogue systems are limited by their reliance on a single dialogue type, which hinders their versatility in practical applications across a range of scenarios, including question-answering, emotional dialog, etc. In this paper, we identify this challenge as how to generate video-driven multilingual mixed-type dialogues. To mitigate this challenge, we propose a novel task and create a human-to-human video-driven multilingual mixed-type dialogue corpus, termed KwaiChat, containing a total of 93,209 videos and 246,080 dialogues, across 4 dialogue types, 30 domains, 4 languages, and 13 topics. Additionally, we establish baseline models on KwaiChat. An extensive analysis of 7 distinct LLMs on KwaiChat reveals that GPT-4o achieves the best performance but still cannot perform well in this situation even with the help of in-context learning and fine-tuning, which indicates that the task is not trivial and needs further research.

cs.CL

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding

With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities of these models. However, these benchmarks focus on standalone videos and mainly assess "visual elements" like human actions and object states. In reality, contemporary videos often encompass complex and continuous narratives, typically presented as a series. To address this challenge, we propose SeriesBench, a benchmark consisting of 105 carefully curated narrative-driven series, covering 28 specialized tasks that require deep narrative understanding. Specifically, we first select a diverse set of drama series spanning various genres. Then, we introduce a novel long-span narrative annotation method, combined with a full-information transformation approach to convert manual annotations into diverse task formats. To further enhance model capacity for detailed analysis of plot structures and character relationships within series, we propose a novel narrative reasoning framework, PC-DCoT. Extensive results on SeriesBench indicate that existing MLLMs still face significant challenges in understanding narrative-driven series, while PC-DCoT enables these MLLMs to achieve performance improvements. Overall, our SeriesBench and PC-DCoT highlight the critical necessity of advancing model capabilities to understand narrative-driven series, guiding the future development of MLLMs. SeriesBench is publicly available at https://github.com/zackhxn/SeriesBench-CVPR2025.

cs.CV

MidMed: Towards Mixed-Type Dialogues for Medical Consultation

Most medical dialogue systems assume that patients have clear goals (medicine querying, surgical operation querying, etc.) before medical consultation. However, in many real scenarios, due to the lack of medical knowledge, it is usually difficult for patients to determine clear goals with all necessary slots. In this paper, we identify this challenge as how to construct medical consultation dialogue systems to help patients clarify their goals. To mitigate this challenge, we propose a novel task and create a human-to-human mixed-type medical consultation dialogue corpus, termed MidMed, covering five dialogue types: task-oriented dialogue for diagnosis, recommendation, knowledge-grounded dialogue, QA, and chitchat. MidMed covers four departments (otorhinolaryngology, ophthalmology, skin, and digestive system), with 8,175 dialogues. Furthermore, we build baselines on MidMed and propose an instruction-guiding medical dialogue generation framework, termed InsMed, to address this task. Experimental results show the effectiveness of InsMed.

cs.CL

AliCHI: A Large-scale Multi-modal Dataset and Automated Evaluation Tool for Human-like Dialogue Systems

A well-designed interactive human-like dialogue system is expected to take actions (e.g. smiling) and respond in a pattern similar to humans. However, due to the limitation of single-modality (only speech) or small volume of currently public datasets, most dialogue systems can only respond in speech and cannot take human-like actions. In this work, we build a large-scale multi-modal dataset of human-to-human conversation in a face-to-face fashion, with fine-grained annotations. The raw data in video format contains 635 dialogue sessions, being collected from 200 participants on designed topics and lasting 52 hours in total. Moreover, we manually annotated the verbal and non-verbal behaviors in each dialogue session on their start/end timestamp. Furthermore, we developed a corresponding evaluation tool for human-like dialogue systems to automatically evaluates the accuracy of two basic tasks, turn-taking prediction, and backchannel prediction, on both time and content. We have opened the data, the tools will be released at the conference.

cs.HC

Residual-type a posteriori error analysis of HDG methods for Neumann boundary control problems

We study a posteriori error analysis of linear-quadratic boundary control problems under bilateral box constraints on the control which acts through a Neumann type boundary condition. We adopt the hybridizable discontinuous Galerkin method as discretization technique, and the flux variables, the scalar variables and the boundary trace variables are all approximated by polynomials of degree k. As for the control variable, it is discretized by the variational discretization concept. Then an efficient and reliable a posteriori error estimator is introduced, and we prove that the error estimator provides an upper bound and a lower bound for the error. Finally, numerical results are presented to illustrate the performance of the obtained a posteriori error estimator.

math.NA

An efficient threshold dynamics method for topology optimization for fluids

We propose an efficient threshold dynamics method for topology optimization for fluids modeled with the Stokes equation. The proposed algorithm is based on minimization of an objective energy function that consists of the dissipation power in the fluid and the perimeter approximated by nonlocal energy, subject to a fluid volume constraint and the incompressibility condition. We show that the minimization problem can be solved with an iterative scheme in which the Stokes equation is approximated by a Brinkman equation. The indicator functions of the fluid-solid regions are then updated according to simple convolutions followed by a thresholding step. We demonstrate mathematically that the iterative algorithm has the total energy decaying property. The proposed algorithm is simple and easy to implement. A simple adaptive time strategy is also used to accelerate the convergence of the iteration. Extensive numerical experiments in both two and three dimensions show that the proposed iteration algorithm converges in much fewer iterations and is more efficient than many existing methods. In addition, the numerical results show that the algorithm is very robust and insensitive to the initial guess and the parameters in the model.

math.OC