SearcharxivSearch

arXiv subjects

Pengzhan Sun

Publications and source records attributed to Pengzhan Sun.

14 recordsLinked to original sources

Dynamic Resolution Routing for Efficient Egocentric Grounding

Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excessive cost of visual token processing. We identify that current efficient strategies based on token reduction are unreliable for selecting object-centric spatial evidence. To overcome this, we propose SmartRes, a framework that performs efficiency optimization in the pixel space via dynamic resolution routing. SmartRes first encodes a low-resolution view for global context and uses a lightweight router to activate high-resolution patches in object-centric regions and constructs an order-preserving visual sequence. To further enable robust routing under severe foreground-background imbalance, we introduce a margin-regularized routing objective that increases foreground-background logit separation and improves foreground recall. Experiments on Ego4D and EgoIntention show that SmartRes reduces visual tokens by up to 67% while retaining 86.4% of full-resolution performance, and achieves up to 1.66X faster inference than state-of-the-art token reduction methods with higher accuracy. Furthermore, strong performance on small object grounding indicates the effectiveness of SmartRes towards egocentric applications. Code will be publicly available.

cs.CV

Ego2World: Compiling Egocentric Cooking Videos into Executable Worlds for Belief-State Planning

Embodied agents in household environments must plan under partial observation: they need to remember objects, track state changes, and recover when actions fail. Existing benchmarks only partially test this ability. Egocentric video datasets capture realistic human activities but remain passive, while interactive simulators support execution but rely on synthetic scenes and hand-crafted dynamics, introducing a sim-to-real gap and often assuming fully observable state. We introduce Ego2World, an executable benchmark that turns egocentric cooking videos into executable symbolic worlds governed by graph-transition rules. Built on HD-EPIC, Ego2World derives reusable transition rules from video annotations and executes them in a hidden symbolic world graph. During evaluation, the simulator maintains the hidden world graph, while the agent plans over its own partial belief graph using only local observations and execution feedback. This separation forces agents to update memory and replan without observing the true world state. Experiments show that action-overlap scores overestimate physical-state success, and that persistent belief memory improves task completion while reducing repeated visual exploration -- suggesting that belief maintenance should be a first-class target of embodied-agent evaluation.

cs.AI

PRISM: : Planning and Reasoning with Intent in Simulated Embodied Environments

When an LLM-based embodied agent fails at a household task, the culprit could be misidentified objects, forgotten sub-goals, or poor action sequencing -- yet existing benchmarks report only a single success rate, making it impossible to tell which cognitive module is responsible. We present PRISM, a diagnostic benchmark that reframes this problem: rather than asking only \textit{did the agent succeed?}, PRISM asks \textit{which capability is most likely responsible for failure?} Built on five photorealistic multi-room apartments (4--8 rooms each), PRISM structures 300 human-verified tasks into three capability tiers -- \textit{Basic Ability}, \textit{Reasoning Ability}, and \textit{Long-horizon Ability} -- that isolate perception-to-action grounding, implicit intent resolution, and sustained multi-step coordination respectively. PRISM exposes an agent-agnostic executable action API that allows arbitrary agents: LLM agents, VLM agents, symbolic planners, RL policies, and hybrid systems, to be evaluated end-to-end under the same benchmark protocol. To support deeper diagnosis, optional probes for perception, memory, and planning can be adopted, replaced, or bypassed entirely, enabling controlled component-level analysis when desired. Experiments on seven contemporary LLMs establish a clear hierarchy: explicit spatial grounding is not the dominant failure source under oracle perception, implicit intent resolution is a significant bottleneck for all model families, and long-horizon coordination exposes a stark capability cliff -- lightweight models collapse to as low as 20.0\% success while simultaneously consuming more tokens than their frontier counterparts, a signature of compensatory over-reasoning rather than genuine planning capability. Project page: \href{https://sj-li.com/PROJ/PRISM}{link}.

cs.RO

Rippled graphene pores as fluidic memristive devices with synaptic and neuromorphic functionalities

Nanofluidic memristive devices work with nanoscale pores and ions dissolved in water, which harness the ionic memory effect aiming to store and process information. These devices share the same charge carriers as biological systems and bring hope for better emulating the neural functions and developing ionic circuits for neuromorphic applications. Specially, theory and experiments suggest that nanoconfinement is essential for inducing a memory effect, which places limit on the pore size to nm-scale or smaller. Such devices are difficult to scale up with precision and operate with long-term stability. Here, we show that a micrometer size pore, generally expected to exhibit a linear ion transport, can display a pronounced memory effect, if its rim is wrapped by strongly curved and tightly stacked graphene. We attribute the observation to slow ion dynamics confined in the rippled graphene edges. The devices are easy to scale up and integrate into fluidic circuits. The memory effect is ion-selective and exhibits long endurance comparable to the lifetime of synaptic proteins, which enables reversible modification of the conductance states using programmable voltage spikes and various electrolytes over a long time, akin to biological synaptic plasticity. Thanks to this plasticity, our devices and their integrated circuits enable storing, transmitting and processing information with high reliability, fidelity and accuracy, as evidenced in the identification of both greyscale and color images, and in the real-time analysis of emulated neural signals. Our results highlight nanoscale morphology of the pore wall as an important parameter regulating ion transport and indicate that the stringent nanoconfinement for ionic memory can be lifted from restricting the pore size to designing its rim structure. The devices and their integrated circuits may find use in ionic neuromorphic applications.

cond-mat.mtrl-sci

Visual Intention Grounding for Egocentric Assistants

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are egocentric, and objects may be referred to implicitly through needs and intentions. To bridge this gap, we introduce EgoIntention, the first dataset for egocentric visual intention grounding. EgoIntention challenges multimodal LLMs to 1) understand and ignore unintended contextual objects and 2) reason about uncommon object functionalities. Benchmark results show that current models misidentify context objects and lack affordance understanding in egocentric views. We also propose Reason-to-Ground (RoG) instruction tuning; it enables hybrid training with normal descriptions and egocentric intentions with a chained intention reasoning and object grounding mechanism. RoG significantly outperforms naive finetuning and hybrid training on EgoIntention, while maintaining or slightly improving naive description grounding. This advancement enables unified visual grounding for egocentric and exocentric visual inputs while handling explicit object queries and implicit human intentions.

cs.CV

Analyzing the Synthetic-to-Real Domain Gap in 3D Hand Pose Estimation

Recent synthetic 3D human datasets for the face, body, and hands have pushed the limits on photorealism. Face recognition and body pose estimation have achieved state-of-the-art performance using synthetic training data alone, but for the hand, there is still a large synthetic-to-real gap. This paper presents the first systematic study of the synthetic-to-real gap of 3D hand pose estimation. We analyze the gap and identify key components such as the forearm, image frequency statistics, hand pose, and object occlusions. To facilitate our analysis, we propose a data synthesis pipeline to synthesize high-quality data. We demonstrate that synthetic hand data can achieve the same level of accuracy as real data when integrating our identified components, paving the path to use synthetic data alone for hand pose estimation. Code and data are available at: https://github.com/delaprada/HandSynthesis.git.

cs.CV

High proton conductivity through angstrom-porous titania

Two dimensional (2D) crystals have attracted strong interest as a new class of proton conducting materials that can block atoms, molecules and ions while allowing proton transport through the atomically thin basal planes. Although 2D materials exhibit this perfect selectivity, the reported proton conductivities have been relatively low. Here we show that vacancy-rich titania monolayers are highly permeable to protons while remaining impermeable to helium with proton conductivity exceeding 100 S cm-2 at 200 C and surpassing targets set by industry roadmaps. The fast and selective proton transport is attributed to an extremely high density of titanium-atom vacancies (one per square nm), which effectively turns titania monolayers into angstrom-scale sieves. Our findings highlight the potential of 2D oxides as membrane materials for hydrogen-based technologies.

cond-mat.mtrl-sci

Simultaneous Detection and Interaction Reasoning for Object-Centric Action Recognition

The interactions between human and objects are important for recognizing object-centric actions. Existing methods usually adopt a two-stage pipeline, where object proposals are first detected using a pretrained detector, and then are fed to an action recognition model for extracting video features and learning the object relations for action recognition. However, since the action prior is unknown in the object detection stage, important objects could be easily overlooked, leading to inferior action recognition performance. In this paper, we propose an end-to-end object-centric action recognition framework that simultaneously performs Detection And Interaction Reasoning in one stage. Particularly, after extracting video features with a base network, we create three modules for concurrent object detection and interaction reasoning. First, a Patch-based Object Decoder generates proposals from video patch tokens. Then, an Interactive Object Refining and Aggregation identifies important objects for action recognition, adjusts proposal scores based on position and appearance, and aggregates object-level info into a global video representation. Lastly, an Object Relation Modeling module encodes object relations. These three modules together with the video feature extractor can be trained jointly in an end-to-end fashion, thus avoiding the heavy reliance on an off-the-shelf object detector, and reducing the multi-stage training burden. We conduct experiments on two datasets, Something-Else and Ikea-Assembly, to evaluate the performance of our proposed approach on conventional, compositional, and few-shot action recognition tasks. Through in-depth experimental analysis, we show the crucial role of interactive objects in learning for action recognition, and we can outperform state-of-the-art methods on both datasets.

cs.CV

HRS-Bench: Holistic, Reliable and Scalable Benchmark for Text-to-Image Models

In recent years, Text-to-Image (T2I) models have been extensively studied, especially with the emergence of diffusion models that achieve state-of-the-art results on T2I synthesis tasks. However, existing benchmarks heavily rely on subjective human evaluation, limiting their ability to holistically assess the model's capabilities. Furthermore, there is a significant gap between efforts in developing new T2I architectures and those in evaluation. To address this, we introduce HRS-Bench, a concrete evaluation benchmark for T2I models that is Holistic, Reliable, and Scalable. Unlike existing bench-marks that focus on limited aspects, HRS-Bench measures 13 skills that can be categorized into five major categories: accuracy, robustness, generalization, fairness, and bias. In addition, HRS-Bench covers 50 scenarios, including fashion, animals, transportation, food, and clothes. We evaluate nine recent large-scale T2I models using metrics that cover a wide range of skills. A human evaluation aligned with 95% of our evaluations on average was conducted to probe the effectiveness of HRS-Bench. Our experiments demonstrate that existing models often struggle to generate images with the desired count of objects, visual text, or grounded emotions. We hope that our benchmark help ease future text-to-image generation research. The code and data are available at https://eslambakr.github.io/hrsbench.github.io

cs.CV

ImageCaptioner$^2$: Image Captioner for Image Captioning Bias Amplification Assessment

Most pre-trained learning systems are known to suffer from bias, which typically emerges from the data, the model, or both. Measuring and quantifying bias and its sources is a challenging task and has been extensively studied in image captioning. Despite the significant effort in this direction, we observed that existing metrics lack consistency in the inclusion of the visual signal. In this paper, we introduce a new bias assessment metric, dubbed $ImageCaptioner^2$, for image captioning. Instead of measuring the absolute bias in the model or the data, $ImageCaptioner^2$ pay more attention to the bias introduced by the model w.r.t the data bias, termed bias amplification. Unlike the existing methods, which only evaluate the image captioning algorithms based on the generated captions only, $ImageCaptioner^2$ incorporates the image while measuring the bias. In addition, we design a formulation for measuring the bias of generated captions as prompt-based image captioning instead of using language classifiers. Finally, we apply our $ImageCaptioner^2$ metric across 11 different image captioning architectures on three different datasets, i.e., MS-COCO caption dataset, Artemis V1, and Artemis V2, and on three different protected attributes, i.e., gender, race, and emotions. Consequently, we verify the effectiveness of our $ImageCaptioner^2$ metric by proposing AnonymousBench, which is a novel human evaluation paradigm for bias metrics. Our metric shows significant superiority over the recent bias metric; LIC, in terms of human alignment, where the correlation scores are 80% and 54% for our metric and LIC, respectively. The code is available at https://eslambakr.github.io/imagecaptioner2.github.io/.

cs.CV

Enhanced hydrogen-gas permeation through rippled graphene

The penetration of atomic hydrogen through defect-free graphene was generally predicted to have a barrier of at least several eV, which is much higher than the 1 eV barrier measured for hydrogen-gas permeation through pristine graphene membranes. Herein, our density functional theory calculations show that ripples, which are ubiquitous in atomically thin crystals and mostly overlooked in the previous simulations, can significantly reduce the barriers for all steps constituting the mechanism of hydrogen-gas permeation through graphene membranes, including dissociation of hydrogen molecules, reconstruction of the dissociated hydrogen atoms and their flipping across graphene. Especially, the flipping barrier of hydrogen atoms from a cluster configuration is found to decrease rapidly down to <1 eV with increasing ripples' curvature. The estimated hydrogen permeation rates by fully considering the distribution of ripples with all realistic curvatures and the major reaction steps that occurred on them are quite close to the experimental measurements. Our work provides insights into the fundamental understanding of hydrogen-gas permeation through graphene membranes and emphasizes the importance of nanoscale non-flatness (ripples) in explaining many surface and transport phenomena (for example, functionalization, corrosion and separation) in graphene and other two-dimensional materials.

cond-mat.mtrl-sci

Highly Efficient and Selective Extraction of Gold by Reduced Graphene Oxide

Materials that are capable of extracting gold from complex sources, especially electronic waste (e-waste) with high efficiency are needed for gold resource sustainability and effective e-waste recycling. However, it remains challenging to achieve high extraction capacity to trace amount of gold, and precise selectivity to gold over a wide range of complex co-existing elements. Here we report a reduced graphene oxide (rGO) material that has an ultrahigh extraction capacity for trace amounts of gold (1,850 mg/g and 1,180 mg/g to 10 ppm and 1 ppm gold). The excellent gold extraction behavior is accounted to the graphene areas and oxidized regions of rGO. The graphene areas spontaneously reduce gold ions to metallic gold, and the oxidized regions provide a good dispersibility so that efficient adsorption and reduction of gold ions by the graphene area can be realized. The rGO is also highly selective to gold ions. By controlling the protonation process of the functional groups on the oxidized regions of rGO, it shows an exclusive gold extraction without adsorption of 14 co-existing elements seen in e-waste. These discoveries are further exploited in highly efficient, continuous gold recycling from e-waste with good scalability and economic viability, as exemplified by extracting gold from e-waste using a rGO membrane based flow-through process.

cond-mat.mtrl-sci

Ultrafast Liquid Water Transport Through Graphene-Based Nanochannels Measured by Isotope Labelling

Graphene-based laminates, with ultralong and tortuous nanocapillaries formed by simply stacking graphene flakes together, have great promises in filtration and separation. However, the information on liquid water trans-membrane permeation is lacking, which is the most fundamental problem and of crucial importance in solution-based mass transport. Here, based on isotope labelling, we investigate the liquid water transportation through graphene-based nanocapillaries under no external hydrostatic pressures. Liquid water can afford an unimpeded permeation through graphene-based nanochannels with a diffusion coefficient 4~5 orders of magnitude larger than through sub-micrometer-sized polymeric channels. When dissolving ions in sources, the diffusion coefficient of ions through graphene channels lies in the same order of magnitude as water, while the ion diffusion is faster than water, indicating that the ions are mainly transported by fast water flows and the delicate interactions between ions and nanocapillary walls also take effect in the accelerated ion transportation.

physics.chem-ph

Structure Evolution of Graphene Oxide during Thermally Driven Phase Transformation: Is the Oxygen Content Really Preserved?

A mild annealing procedure was recently proposed for the scalable enhancement of graphene oxide (GO) properties with the oxygen content preserved, which was demonstrated to be attributed to the thermally driven phase separation. In this work, the structure evolution of GO with mild annealing is closely investigated. It reveals that in addition to phase separation, the transformation of oxygen functionalities also occurs, which leads to the slight reduction of GO membranes and furthers the enhancement of GO properties. These results are further supported by the density functional theory based calculations. The results also show that the amount of chemically bonded oxygen atoms on graphene decreases gradually and we propose that the strongly physisorbed oxygen species constrained in the holes and vacancies on GO lattice might be responsible for the preserved oxygen content during the mild annealing procedure. The present experimental results and calculations indicate that both the diffusion and transformation of oxygen functional groups might play important roles in the scalable enhancement of GO properties.

physics.chem-ph