Searcharxiv⌕ Search

arXiv subjects

Ping Hu

Publications and source records attributed to Ping Hu.

At least 55 records · Page 3Linked to original sources

Deepsea: A Meta-ocean Prototype for Undersea Exploration

Metaverse has attracted great attention from industry and academia in recent years. Metaverse for the ocean (Meta-ocean) is the implementation of the Metaverse technologies in virtual emersion of the ocean which is beneficial for people yearning for the ocean. It has demonstrated great potential for tourism and education with its strong immersion and appealing interactive user experience. However, quite limited endeavors have been spent on exploring the full possibility of Meta-ocean, especially in modeling the movements of marine creatures. In this paper, we first investigate the technology status of Metaverse and virtual reality (VR) and develop a prototype that builds the Meta-ocean in VR devices with strong immersive visual effects. Then, we demonstrate a method to model the undersea scene and marine creatures and propose an optimized path algorithm based on the Catmull-Rom spline to model the movements of marine life. Finally, we conduct a user study to analyze our Meta-ocean prototype. This user study illustrates that our new prototype can give us strong immersion and an appealing interactive user experience.

cs.HC↗

Optimal Aggregation Strategies for Social Learning over Graphs

Adaptive social learning is a useful tool for studying distributed decision-making problems over graphs. This paper investigates the effect of combination policies on the performance of adaptive social learning strategies. Using large-deviation analysis, it first derives a bound on the steady-state error probability and characterizes the optimal selection for the Perron eigenvectors of the combination policies. It subsequently studies the effect of the combination policy on the transient behavior of the learning strategy by estimating the adaptation time in the low signal-to-noise ratio regime. In the process, it is discovered that, interestingly, the influence of the combination policy on the transient behavior is insignificant, and thus it is more critical to employ policies that enhance the steady-state performance. The theoretical conclusions are illustrated by means of computer simulations.

eess.SP↗

Token Boosting for Robust Self-Supervised Visual Transformer Pre-training

Learning with large-scale unlabeled data has become a powerful tool for pre-training Visual Transformers (VTs). However, prior works tend to overlook that, in real-world scenarios, the input data may be corrupted and unreliable. Pre-training VTs on such corrupted data can be challenging, especially when we pre-train via the masked autoencoding approach, where both the inputs and masked ``ground truth" targets can potentially be unreliable in this case. To address this limitation, we introduce the Token Boosting Module (TBM) as a plug-and-play component for VTs that effectively allows the VT to learn to extract clean and robust features during masked autoencoding pre-training. We provide theoretical analysis to show how TBM improves model pre-training with more robust and generalizable representations, thus benefiting downstream tasks. We conduct extensive experiments to analyze TBM's effectiveness, and results on four corrupted datasets demonstrate that TBM consistently improves performance on downstream tasks.

cs.CV↗

Clique-factors in graphs with sublinear $\ell$-independence number

Given a graph $G$ and an integer $\ell\ge 2$, we denote by $α_{\ell}(G)$ the maximum size of a $K_{\ell}$-free subset of vertices in $V(G)$. A recent question of Nenadov and Pehova asks for determining the best possible minimum degree conditions forcing clique-factors in $n$-vertex graphs $G$ with $α_{\ell}(G) = o(n)$, which can be seen as a Ramsey--Turán variant of the celebrated Hajnal--Szemerédi theorem. In this paper we find the asymptotical sharp minimum degree threshold for $K_r$-factors in $n$-vertex graphs $G$ with $α_\ell(G)=n^{1-o(1)}$ for all $r\ge \ell\ge 2$.

math.CO↗

Splatting-based Synthesis for Video Frame Interpolation

Frame interpolation is an essential video processing technique that adjusts the temporal resolution of an image sequence. While deep learning has brought great improvements to the area of video frame interpolation, techniques that make use of neural networks can typically not easily be deployed in practical applications like a video editor since they are either computationally too demanding or fail at high resolutions. In contrast, we propose a deep learning approach that solely relies on splatting to synthesize interpolated frames. This splatting-based synthesis for video frame interpolation is not only much faster than similar approaches, especially for multi-frame interpolation, but can also yield new state-of-the-art results at high resolutions.

cs.CV↗

DualCoOp: Fast Adaptation to Multi-Label Recognition with Limited Annotations

Solving multi-label recognition (MLR) for images in the low-label regime is a challenging task with many real-world applications. Recent work learns an alignment between textual and visual spaces to compensate for insufficient image labels, but loses accuracy because of the limited amount of available MLR annotations. In this work, we utilize the strong alignment of textual and visual features pretrained with millions of auxiliary image-text pairs and propose Dual Context Optimization (DualCoOp) as a unified framework for partial-label MLR and zero-shot MLR. DualCoOp encodes positive and negative contexts with class names as part of the linguistic input (i.e. prompts). Since DualCoOp only introduces a very light learnable overhead upon the pretrained vision-language framework, it can quickly adapt to multi-label recognition tasks that have limited annotations and even unseen classes. Experiments on standard multi-label recognition benchmarks across two challenging low-label settings demonstrate the advantages of our approach over state-of-the-art methods.

cs.CV↗

ZeroWaste Dataset: Towards Deformable Object Segmentation in Cluttered Scenes

Less than 35% of recyclable waste is being actually recycled in the US, which leads to increased soil and sea pollution and is one of the major concerns of environmental researchers as well as the common public. At the heart of the problem are the inefficiencies of the waste sorting process (separating paper, plastic, metal, glass, etc.) due to the extremely complex and cluttered nature of the waste stream. Recyclable waste detection poses a unique computer vision challenge as it requires detection of highly deformable and often translucent objects in cluttered scenes without the kind of context information usually present in human-centric datasets. This challenging computer vision task currently lacks suitable datasets or methods in the available literature. In this paper, we take a step towards computer-aided waste detection and present the first in-the-wild industrial-grade waste detection and segmentation dataset, ZeroWaste. We believe that ZeroWaste will catalyze research in object detection and semantic segmentation in extreme clutter as well as applications in the recycling domain. Our project page can be found at http://ai.bu.edu/zerowaste/.

cs.CV↗

Learning to Detect Every Thing in an Open World

Many open-world applications require the detection of novel objects, yet state-of-the-art object detection and instance segmentation networks do not excel at this task. The key issue lies in their assumption that regions without any annotations should be suppressed as negatives, which teaches the model to treat the unannotated objects as background. To address this issue, we propose a simple yet surprisingly powerful data augmentation and training scheme we call Learning to Detect Every Thing (LDET). To avoid suppressing hidden objects, background objects that are visible but unlabeled, we paste annotated objects on a background image sampled from a small region of the original image. Since training solely on such synthetically-augmented images suffers from domain shift, we decouple the training into two parts: 1) training the region classification and regression head on augmented images, and 2)~training the mask heads on original images. In this way, a model does not learn to classify hidden objects as background while generalizing well to real images. LDET leads to significant improvements on many datasets in the open-world instance segmentation task, outperforming baselines on cross-category generalization on COCO, as well as cross-dataset evaluation on UVO and Cityscapes.

cs.CV↗

Many-to-many Splatting for Efficient Video Frame Interpolation

Motion-based video frame interpolation commonly relies on optical flow to warp pixels from the inputs to the desired interpolation instant. Yet due to the inherent challenges of motion estimation (e.g. occlusions and discontinuities), most state-of-the-art interpolation approaches require subsequent refinement of the warped result to generate satisfying outputs, which drastically decreases the efficiency for multi-frame interpolation. In this work, we propose a fully differentiable Many-to-Many (M2M) splatting framework to interpolate frames efficiently. Specifically, given a frame pair, we estimate multiple bidirectional flows to directly forward warp the pixels to the desired time step, and then fuse any overlapping pixels. In doing so, each source pixel renders multiple target pixels and each target pixel can be synthesized from a larger area of visual context. This establishes a many-to-many splatting scheme with robustness to artifacts like holes. Moreover, for each input frame pair, M2M only performs motion estimation once and has a minuscule computational overhead when interpolating an arbitrary number of in-between frames, hence achieving fast multi-frame interpolation. We conducted extensive experiments to analyze M2M, and found that it significantly improves efficiency while maintaining high effectiveness.

cs.CV↗

Geometry-Aware Planar Embedding of Treelike Structures

The growing complexity of spatial and structural information in 3D data makes data inspection and visualization a challenging task. We describe a method to create a planar embedding of 3D treelike structures using their skeleton representations. Our method maintains the original geometry, without overlaps, to the best extent possible, allowing exploration of the topology within a single view. We present a novel camera view generation method which maximizes the visible geometric attributes (segment shape and relative placement between segments). Camera views are created for individual segments and are used to determine local bending angles at each node by projecting them to 2D. The final embedding is generated by minimizing an energy function (the weights of which are user adjustable) based on branch length and the 2D angles, while avoiding intersections. The user can also interactively modify segment placement within the 2D embedding, and the overall embedding will update accordingly. A global to local interactive exploration is provided using hierarchical camera views that are created for subtrees within the structure. We evaluate our method both qualitatively and quantitatively and demonstrate our results by constructing planar visualizations of line data (traced neurons) and volume data (CT vascular and bronchial data

cs.GR↗

Tilings in graphons

We introduce a counterpart to the notion of vertex disjoint tilings by copy of a fixed graph F to the setting of graphons. The case F=K_2 gives the notion of matchings in graphons. We give a transference statement that allows us to switch between the finite and limit notion, and derive several favorable properties, including the LP-duality counterpart to the classical relation between the fractional vertex covers and fractional matchings/tilings, and discuss connections with property testing. As an application of our theory, we determine the asymptotically almost sure F-tiling number of inhomogeneous random graphs \mathbb{G}(n,W). As another application, in an accompanying paper [Hladky, Hu, Piguet: Komlos's tiling theorem via graphon covers, preprint] we give a proof of a strengthening of a theorem of Komlos [Komlos: Tiling Turán Theorems, Combinatorica, 2000].

math.CO↗

The inducibility of oriented stars

We consider the problem of maximizing the number of induced copies of an oriented star $S_{k,\ell}$ in digraphs of given size, where the center of the star has out-degree $k$ and in-degree $\ell$. The case $k\ell=0$ was solved by Huang. Here, we asymptotically solve it for all other oriented stars with at least seven vertices.

math.CO↗

Real-time Semantic Segmentation with Fast Attention

In deep CNN based models for semantic segmentation, high accuracy relies on rich spatial context (large receptive fields) and fine spatial details (high resolution), both of which incur high computational costs. In this paper, we propose a novel architecture that addresses both challenges and achieves state-of-the-art performance for semantic segmentation of high-resolution images and videos in real-time. The proposed architecture relies on our fast spatial attention, which is a simple yet efficient modification of the popular self-attention mechanism and captures the same rich spatial context at a small fraction of the computational cost, by changing the order of operations. Moreover, to efficiently process high-resolution input, we apply an additional spatial reduction to intermediate feature stages of the network with minimal loss in accuracy thanks to the use of the fast attention module to fuse features. We validate our method with a series of experiments, and show that results on multiple datasets demonstrate superior performance with better accuracy and speed compared to existing approaches for real-time semantic segmentation. On Cityscapes, our network achieves 74.4$\%$ mIoU at 72 FPS and 75.5$\%$ mIoU at 58 FPS on a single Titan X GPU, which is~$\sim$50$\%$ faster than the state-of-the-art while retaining the same accuracy.

cs.CV↗

Temporally Distributed Networks for Fast Video Semantic Segmentation

We present TDNet, a temporally distributed network designed for fast and accurate video semantic segmentation. We observe that features extracted from a certain high-level layer of a deep CNN can be approximated by composing features extracted from several shallower sub-networks. Leveraging the inherent temporal continuity in videos, we distribute these sub-networks over sequential frames. Therefore, at each time step, we only need to perform a lightweight computation to extract a sub-features group from a single sub-network. The full features used for segmentation are then recomposed by application of a novel attention propagation module that compensates for geometry deformation between frames. A grouped knowledge distillation loss is also introduced to further improve the representation power at both full and sub-feature levels. Experiments on Cityscapes, CamVid, and NYUD-v2 demonstrate that our method achieves state-of-the-art accuracy with significantly faster speed and lower latency.

cs.CV↗

Weakly-supervised Compositional FeatureAggregation for Few-shot Recognition

Learning from a few examples is a challenging task for machine learning. While recent progress has been made for this problem, most of the existing methods ignore the compositionality in visual concept representation (e.g. objects are built from parts or composed of semantic attributes), which is key to the human ability to easily learn from a small number of examples. To enhance the few-shot learning models with compositionality, in this paper we present the simple yet powerful Compositional Feature Aggregation (CFA) module as a weakly-supervised regularization for deep networks. Given the deep feature maps extracted from the input, our CFA module first disentangles the feature space into disjoint semantic subspaces that model different attributes, and then bilinearly aggregates the local features within each of these subspaces. CFA explicitly regularizes the representation with both semantic and spatial compositionality to produce discriminative representations for few-shot recognition tasks. Moreover, our method does not need any supervision for attributes and object parts during training, thus can be conveniently plugged into existing models for end-to-end optimization while keeping the model size and computation cost nearly the same. Extensive experiments on few-shot image classification and action recognition tasks demonstrate that our method provides substantial improvements over recent state-of-the-art methods.

cs.CV↗

Minimum number of edges that occur in odd cycles

If a graph has $n\ge4k$ vertices and more than $n^2/4$ edges, then it contains a copy of $C_{2k+1}$. In 1992, Erdős, Faudree and Rousseau showed even more, that the number of edges that occur in a triangle is at least $2\lfloor n/2\rfloor+1$, and this bound is tight. They also showed that the minimum number of edges that occur in a $C_{2k+1}$ for $k\ge2$ is at least $11n^2/144-O(n)$, and conjectured that for any $k\ge2$, the correct lower bound should be $2n^2/9-O(n)$. Very recently, Füredi and Maleki constructed a counterexample for $k=2$ and proved asymptotically matching lower bound, namely that for any $\varepsilon>0$ graphs with $(1+\varepsilon)n^2/4$ edges contain at least $(2+\sqrt{2})n^2/16 \approx 0.2134n^2$ edges that occur in $C_5$. In this paper, we use a different approach to tackle this problem and obtain the following stronger result: Any $n$-vertex graph with at least $\lfloor n^2/4\rfloor+1$ edges has at least $(2+\sqrt{2})n^2/16-O(n^{15/8})$ edges that occur in $C_5$. Next, for all $k\ge 3$ and $n$ sufficiently large, we determine the exact minimum number of edges that occur in $C_{2k+1}$ for $n$-vertex graphs with more than $n^2/4$ edges, and show it is indeed equal to $\lfloor\frac{n^2}4\rfloor+1-\lfloor\frac{n+4}6\rfloor\lfloor\frac{n+1}6\rfloor=2n^2/9-O(n)$. For both results, we give a structural description of the extremal configurations as well as obtain the corresponding stability results, which answer a conjecture of Füredi and Maleki. The main ingredient is a novel approach that combines the flag algebras together with ideas from finite forcibility of graph limits. This approach allowed us to keep track of the extra edge needed to guarantee an existence of a $C_{2k+1}$. Also, we establish the first application of semidefinite method in a setting, where the set of tight examples has exponential size, and arises from different constructions.

math.CO↗

Skew-symmetric Nitsche's formulation in isogeometric analysis: Dirichlet and symmetry conditions, patch coupling and frictionless contact

A simple skew-symmetric Nitsche's formulation is introduced into the framework of isogeometric analysis (IGA) to deal with various problems in small strain elasticity: essential boundary conditions, symmetry conditions for Kirchhoff plates, patch coupling in statics and in modal analysis as well as Signorini contact conditions. For linear boundary or interface conditions, the skew-symmetric formulation is parameter-free. For contact conditions, it remains stable and accurate for a wide range of the stabilization parameter. Several numerical tests are performed to illustrate its accuracy, stability and convergence performance. We investigate particularly the effects introduced by Nitsche's coupling, including the convergence performance and condition numbers in statics as well as the extra "outlier" frequencies and corresponding eigenmodes in structural dynamics. We present the Hertz test, the block test, and a 3D self-contact example showing that the skew-symmetric Nitsche's formulation is a suitable approach to simulate contact problems in IGA.

math.NA↗

Komlós's tiling theorem via graphon covers

Komlos [Komlos: Tiling Turan Theorems, Combinatorica, 2000] determined the asymptotically optimal minimum-degree condition for covering a given proportion of vertices of a host graph by vertex-disjoint copies of a fixed graph H, thus essentially extending the Hajnal-Szemeredi theorem which deals with the case when H is a clique. We give a proof of a graphon version of Komlos's theorem. To prove this graphon version, and also to deduce from it the original statement about finite graphs, we use the machinery introduced in [Hladky, Hu, Piguet: Tilings in graphons, arXiv:1606.03113]. We further prove a stability version of Komlos's theorem.

math.CO↗