SearcharxivSearch

arXiv subjects

Lu He

Publications and source records attributed to Lu He.

16 recordsLinked to original sources

PolycubeNet: A Dual-latent Diffusion Model for Polycube-Based Hexahedral Mesh Generation

Hexahedral meshes are widely used in simulation pipelines, yet automatic generation remains challenging for complex CAD geometries. Polycube-based hexahedral meshing is a representative approach due to its regular, parameterization-friendly structure, but existing polycube construction methods often rely on intricate surface segmentation and local heuristics, which can produce artifacts or fail on difficult shapes. In this paper, we propose an end-to-end framework for polycube generation based on conditional diffusion models. Given an input geometry represented as a point cloud, our method directly produces a corresponding polycube point cloud, eliminating the need for explicit surface segmentation or predefined polycube templates. At the core of our approach is a dual-latent conditional diffusion architecture that confines computationally expensive self-attention operations to a fixed-capacity, low-dimensional latent space. This design effectively decouples computational complexity from the resolution of both the input geometry and the output polycube, thereby avoiding the quadratic cost typical of point cloud self-attention mechanisms while supporting flexible input and output resolutions. To obtain a hexahedral mesh, the generated polycube is aligned to the input shape via rigid and non-rigid point cloud registration to establish surface correspondence, followed by a polycube-to-hex pipeline. We additionally create and release a paired dataset of CAD meshes and their corresponding polycube meshes, together with the core implementation of our model. Experiments show that PolycubeNet generalizes to complex CAD models with arbitrary genus and produces high-quality polycube structures within seconds, improving robustness and efficiency over prior learning-based approaches.

cs.GR

Activation Function Optimization Scheme for Image Classification

Activation function has a significant impact on the dynamics, convergence, and performance of deep neural networks. The search for a consistent and high-performing activation function has always been a pursuit during deep learning model development. Existing state-of-the-art activation functions are manually designed with human expertise except for Swish. Swish was developed using a reinforcement learning-based search strategy. In this study, we propose an evolutionary approach for optimizing activation functions specifically for image classification tasks, aiming to discover functions that outperform current state-of-the-art options. Through this optimization framework, we obtain a series of high-performing activation functions denoted as Exponential Error Linear Unit (EELU). The developed activation functions are evaluated for image classification tasks from two perspectives: (1) five state-of-the-art neural network architectures, such as ResNet50, AlexNet, VGG16, MobileNet, and Compact Convolutional Transformer which cover computationally heavy to light neural networks, and (2) eight standard datasets, including CIFAR10, Imagenette, MNIST, Fashion MNIST, Beans, Colorectal Histology, CottonWeedID15, and TinyImageNet which cover from typical machine vision benchmark, agricultural image applications to medical image applications. Finally, we statistically investigate the generalization of the resultant activation functions developed through the optimization scheme. With a Friedman test, we conclude that the optimization scheme is able to generate activation functions that outperform the existing standard ones in 92.8% cases among 28 different cases studied, and $-x\cdot erf(e^{-x})$ is found to be the best activation function for image classification generated by the optimization scheme.

cs.CV

Wonderful Matrices: More Efficient and Effective Architecture for Language Modeling Tasks

We prove the availability of inner product form position encoding in the state space dual algorithm and study the effectiveness of different position embeddings in the hybrid quadratic causal self-attention and state space dual algorithms. We propose inner function attention with dynamic mask, which can improve the expressiveness of the attention algorithm and avoid the sequence noise significantly affecting the accuracy of the attention score. We also design cross domain mixture of experts, which can improve the granularity of the sparse activation feedforward network while maintaining the efficiency of parameter utilization and retrieval. The combination of these methods constitutes our foundation model architecture: Wonderful Matrices. We conduct experiments on the language modeling task and find that Wonderful Matrices are more efficient and effective in handling complex language tasks.

cs.LG

Hyperbolic photonic topological insulators

Topological photonics provides a new degree of freedom to robustly control electromagnetic fields. To date, most of established topological states in photonics have been employed in Euclidean space. Motivated by unique properties of hyperbolic lattices, which are regular tessellations in non-Euclidean space with a constant negative curvature, the boundarydominated hyperbolic topological states have been proposed. However, limited by highly crowded boundary resonators and complicated site couplings, the hyperbolic topological insulator has only been experimentally constructed in electric circuits. How to achieve hyperbolic photonic topological insulators is still an open question. Here, we report the experimental realization of hyperbolic photonic topological insulators using coupled ring resonators on silicon chips. Boundary-dominated one-way edge states with pseudospindependent propagation directions have been observed. Furthermore, the robustness of edge states in hyperbolic photonic topological insulators is also verified. Our findings have potential applications in the field of designing high-efficient topological photonic devices with enhanced boundary responses.

physics.optics

Engage Wider Audience or Facilitate Quality Answers? a Mixed-methods Analysis of Questioning Strategies for Research Sensemaking on a Community Q&A Site

Discussing research-sensemaking questions on Community Question and Answering (CQA) platforms has been an increasingly common practice for the public to participate in science communication. Nonetheless, how users strategically craft research-sensemaking questions to engage public participation and facilitate knowledge construction is a significant yet less understood problem. To fill this gap, we collected 837 science-related questions and 157,684 answers from Zhihu, and conducted a mixed-methods study to explore user-developed strategies in proposing research-sensemaking questions, and their potential effects on public engagement and knowledge construction. Through open coding, we captured a comprehensive taxonomy of question-crafting strategies, such as eyecatching narratives with counter-intuitive claims and rigorous descriptions with data use. Regression analysis indicated that these strategies correlated with user engagement and answer construction in different ways (e.g., emotional questions attracted more views and answers), yet there existed a general divergence between wide participation and quality knowledge establishment, when most questioning strategies could not ensure both. Based on log analysis, we further found that collaborative editing afforded unique values in refining research-sensemaking questions regarding accuracy, rigor, comprehensiveness and attractiveness. We propose design implications to facilitate accessible, accurate and engaging science communication on CQA platforms.

cs.HC

Images Connect Us Together: Navigating a COVID-19 Local Outbreak in China Through Social Media Images

Social media images, curated or casual, have become a crucial component of communicating situational information and emotions during health crises. Despite its prevalence and significance in informational dissemination and emotional connection, there lacks a comprehensive understanding of visual crisis communication in the aftermath of a pandemic which is characterized by uncertain local situations and emotional fatigue. To fill this gap, this work collected 345,423 crisis-related posts and 65,376 original images during the Xi'an COVID-19 local outbreak in China, and adopted a mixed-methods approach to understanding themes, goals, and strategies of crisis imagery. Image clustering captured the diversity of visual themes during the outbreak, such as text images embedding authoritative guidelines and ``visual diaries'' recording and sharing the quarantine life. Through text classification of the post that visuals were situated in, we found that different visual themes highly correlated with the informational and emotional goals of the post text, such as adopting text images to convey the latest policies and sharing food images to express anxiety. We further unpacked nuanced strategies of crisis image use through inductive coding, such as signifying authority and triggering empathy. We discuss the opportunities and challenges of crisis imagery and provide design implications to facilitate effective visual crisis communication.

cs.HC

Super-compact universal quantum logic gates with inversedesigned elements

Integrated quantum photonic circuit is a promising platform for the realization of quantum information processing in the future. To achieve the largescale quantum photonic circuits, the applied quantum logic gates should be as small as possible for the high-density integration on chips. Here, we report the implementation of super-compact universal quantum logic gates on silicon chips by the method of inverse design. In particular, the fabricated controlled-NOT gate and Hadamard gate are both nearly a vacuum wavelength, being the smallest optical quantum gates reported up to now. We further design the quantum circuit by cascading these fundamental gates to perform arbitrary quantum processing, where the corresponding size is about several orders smaller than that of previous quantum photonic circuits. Our study paves the way for the realization of largescale quantum photonic chips with integrated sources, and can possess important applications in the field of quantum information processes.

physics.optics

Experimental realization of topologically-protected all-optical logic gates based on silicon photonic crystal slabs

Topological photonics has been developed for more than ten years. It has been proved that the combination of topology and photons is very beneficial to the design of robust optical devices against some disturbances. However, most of the work for robust optical logic devices stays at the theoretical level. There are very few topologically-protected logic devices fabricated in experiments. Here, we report the experimental fabrication of a series of topologically-protected all-optical logic gates. Seven topologically-protected all-optical logic gates (OR, XOR, NOT, XNOR, NAND, NOR, and AND) are fabricated on silicon photonic platforms, which show strong robustness even if some disorders exist. These robust logic devices are potentially applicable in future optical signal processing and computing.

physics.optics

PanelNet: Understanding 360 Indoor Environment via Panel Representation

Indoor 360 panoramas have two essential properties. (1) The panoramas are continuous and seamless in the horizontal direction. (2) Gravity plays an important role in indoor environment design. By leveraging these properties, we present PanelNet, a framework that understands indoor environments using a novel panel representation of 360 images. We represent an equirectangular projection (ERP) as consecutive vertical panels with corresponding 3D panel geometry. To reduce the negative impact of panoramic distortion, we incorporate a panel geometry embedding network that encodes both the local and global geometric features of a panel. To capture the geometric context in room design, we introduce Local2Global Transformer, which aggregates local information within a panel and panel-wise global context. It greatly improves the model performance with low training overhead. Our method outperforms existing methods on indoor 360 depth estimation and shows competitive results against state-of-the-art approaches on the task of indoor layout estimation and semantic segmentation.

cs.CV

Weakly-Supervised Temporal Action Detection for Fine-Grained Videos with Hierarchical Atomic Actions

Action understanding has evolved into the era of fine granularity, as most human behaviors in real life have only minor differences. To detect these fine-grained actions accurately in a label-efficient way, we tackle the problem of weakly-supervised fine-grained temporal action detection in videos for the first time. Without the careful design to capture subtle differences between fine-grained actions, previous weakly-supervised models for general action detection cannot perform well in the fine-grained setting. We propose to model actions as the combinations of reusable atomic actions which are automatically discovered from data through self-supervised clustering, in order to capture the commonality and individuality of fine-grained actions. The learnt atomic actions, represented by visual concepts, are further mapped to fine and coarse action labels leveraging the semantic label hierarchy. Our approach constructs a visual representation hierarchy of four levels: clip level, atomic action level, fine action class level and coarse action class level, with supervision at each level. Extensive experiments on two large-scale fine-grained video datasets, FineAction and FineGym, show the benefit of our proposed weakly-supervised model for fine-grained action detection, and it achieves state-of-the-art results.

cs.CV

Exciton tuning in monolayer WSe$_2$ via substrate induced electron doping

We report on large exciton tuning in WSe$_2$ monolayers via substrate induced non-degenerate doping. We observe a redshift of $\sim$62 meV for the $A$ exciton together with a 1-2 orders of magnitude photoluminescence (PL) quenching when the monolayer WSe$_2$ is brought in contact with highly oriented pyrolytic graphite (HOPG) compared to the dielectric substrates such as hBN and SiO$_2$. As the evidence of doping from HOPG to WSe$_2$, a drastic increase of the trion emission intensity was observed. Using a systematic PL and Kelvin probe force microscopy (KPFM) investigation on WSe$_2$/HOPG, WSe$_2$/hBN, and WSe$_2$/graphene, we conclude that this unique excitonic behavior is induced by electron doping from the substrate. Our results propose a simple yet efficient way for exciton tuning in monolayer WSe$_2$, which plays a central role in the fundamental understanding and further device development.

cond-mat.mtrl-sci

TransVOD: End-to-End Video Object Detection with Spatial-Temporal Transformers

Detection Transformer (DETR) and Deformable DETR have been proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance as previous complex hand-crafted detectors. However, their performance on Video Object Detection (VOD) has not been well explored. In this paper, we present TransVOD, the first end-to-end video object detection system based on spatial-temporal Transformer architectures. The first goal of this paper is to streamline the pipeline of VOD, effectively removing the need for many hand-crafted components for feature aggregation, e.g., optical flow model, relation networks. Besides, benefited from the object query design in DETR, our method does not need complicated post-processing methods such as Seq-NMS. In particular, we present a temporal Transformer to aggregate both the spatial object queries and the feature memories of each frame. Our temporal transformer consists of two components: Temporal Query Encoder (TQE) to fuse object queries, and Temporal Deformable Transformer Decoder (TDTD) to obtain current frame detection results. These designs boost the strong baseline deformable DETR by a significant margin (3%-4% mAP) on the ImageNet VID dataset. Then, we present two improved versions of TransVOD including TransVOD++ and TransVOD Lite. The former fuses object-level information into object query via dynamic convolution while the latter models the entire video clips as the output to speed up the inference time. We give detailed analysis of all three models in the experiment part. In particular, our proposed TransVOD++ sets a new state-of-the-art record in terms of accuracy on ImageNet VID with 90.0% mAP. Our proposed TransVOD Lite also achieves the best speed and accuracy trade-off with 83.7% mAP while running at around 30 FPS on a single V100 GPU device.

cs.CV

Topology-optimized ultra-compact all-optical logic devices on silicon photonic platforms

The realization of all-optical integration and optical computing has always been our goal. One of the most significant challenges is to make integrated all-optical logic devices as small as possible. Here, we report the implementation of ultra-compact all-optical logic devices and integrated chips on silicon photonic platforms by topology optimization. The footprint for the fabricated all-optical logic gates with XOR and OR functions is only 1.3*1.3 {\mu}m2 (~0.84{\lambda}*0.84{\lambda}), that are the smallest all-optical dielectric logic devices ever verified in experiments in the optical communication range. The ultra-low loss of the optical signal is also demonstrated experimentally (-0.96dB). Furthermore, an integrated chip containing seven major logic gates (AND, OR, NOT, NAND, NOR, XOR, and XNOR) and a half adder is fabricated, where the associated footprint is only 1.3*4.5 {\mu}m2. Our work opens up a new path towards practical all-optical integration and optical computing.

physics.optics

End-to-End Video Object Detection with Spatial-Temporal Transformers

Recently, DETR and Deformable DETR have been proposed to eliminate the need for many hand-designed components in object detection while demonstrating good performance as previous complex hand-crafted detectors. However, their performance on Video Object Detection (VOD) has not been well explored. In this paper, we present TransVOD, an end-to-end video object detection model based on a spatial-temporal Transformer architecture. The goal of this paper is to streamline the pipeline of VOD, effectively removing the need for many hand-crafted components for feature aggregation, e.g., optical flow, recurrent neural networks, relation networks. Besides, benefited from the object query design in DETR, our method does not need complicated post-processing methods such as Seq-NMS or Tubelet rescoring, which keeps the pipeline simple and clean. In particular, we present temporal Transformer to aggregate both the spatial object queries and the feature memories of each frame. Our temporal Transformer consists of three components: Temporal Deformable Transformer Encoder (TDTE) to encode the multiple frame spatial details, Temporal Query Encoder (TQE) to fuse object queries, and Temporal Deformable Transformer Decoder to obtain current frame detection results. These designs boost the strong baseline deformable DETR by a significant margin (3%-4% mAP) on the ImageNet VID dataset. TransVOD yields comparable results performance on the benchmark of ImageNet VID. We hope our TransVOD can provide a new perspective for video object detection. Code will be made publicly available at https://github.com/SJTU-LuHe/TransVOD.

cs.CV

Topologically protected long-range coherent energy transfer

The realization of robust coherent energy transfer with a long range from a donor to an acceptor has many important applications in the field of quantum optics. However, it is hard to be realized using conventional schemes. Here, we demonstrate theoretically that the robust energy transfer can be achieved using a photonic crystal platform, which includes the topologically protected edge state and zero-dimensional topological corner cavities. When the donor and the acceptor are put into a pair of separated topological cavities, respectively, the energy transfer between them can be fulfilled with the assistance of the topologically protected interface state. Such an energy transfer is robust against various kinds of defects, and can also occur over very long distances, which is very beneficial for biological detections, sensors, quantum information science and so on.

physics.optics

Topologically protected strong coupling and entanglement between distant quantum emitters

The realization of robust strong coupling and entanglement between distant quantum emitters (QEs) is very important for scalable quantum information processes. However, it is hard to achieve it based on conventional systems. Here, we propose theoretically and demonstrate numerically a scheme to realize such strong coupling and entanglement. Our scheme is based on the photonic crystal platform with topologically protected edge state and zero-dimensional topological corner cavities. When the QEs are put into topological cavities, the strong coupling between them can be fulfilled with the assistance of the topologically protected interface state. Such a strong coupling can maintain a very long distance and be robust against various defects. Especially, we numerically prove that the topologically protected entanglement between two QEs can also be realized. Moreover, the duration of quantum beats for such entanglement can reach several orders longer than that for the entanglement in a conventional photonic cavity, making it be very beneficial for a scalable quantum information process.

physics.optics