SearcharxivSearch

arXiv subjects

Nikola Popovic

Publications and source records attributed to Nikola Popovic.

17 recordsLinked to original sources

Cell Division Changes Fate Decisions in a Genetic Toggle Switch

Gene regulatory networks govern cellular fate decisions through multistable dynamics. The genetic toggle switch is a canonical model of such behaviour; yet, the impact of cell division on its dynamics remains poorly understood. We derive analytical separatrices for a simplified Boolean toggle switch with and without division. We show that division can redirect trajectories with identical initial conditions to opposing stable states, and we define a region of disagreement where fate decisions are predicted incorrectly if division is neglected. Our results imply that division can fundamentally reshape fate boundaries in multistable regulatory networks.

q-bio.MN

Rate-induced tipping in a coral reef ecosystem: A slow increase in fishing effort can induce reef collapse

Critical transitions describe sudden changes in the state of an ecosystem. In classical bifurcation theory, such transitions occur when the value of a parameter exceeds a threshold (``bifurcation") value. More recently, critical transitions which are triggered by the rate of change of a parameter were described by Wieczorek et al. [Wieczorek, S., Ashwin, P., Luke, C.M., Cox, P.M., Proceedings of the Royal Society A 467(2129), 1243-1269, 2011]. In mathematical ecology, these rate-induced transitions correspond to environmental conditions that deteriorate too rapidly for the ecosystem to adapt, resulting in population collapse (``R-tipping"). In this article, we consider the potential for rate-induced tipping due to increased anthropogenic stress in a recently proposed behavioural-demographic model for herbivorous fish, algae, and coral in a coral reef ecosystem [Gil, M.A., Baskett, M.L., Munch, S.B., Hein, A.M., PNAS 117(41), 25580-25589, 2020]. We first show that the underlying demographic model can be reframed naturally as a singularly perturbed system with two fast variables and one slow variable in which bistability can occur in ecologically relevant parameter regimes. We explore the potential for canard-type dynamics in the model, complementing numerical results with an analytical description through the lens of geometric singular perturbation theory, and we describe R-tipping as a result of an increase in the fishing effort. We show that trajectories will undergo canard-induced tipping by passage through a folded node singularity, whereas a folded focus may give rise to tipping of jump type; in both scenarios, a catastrophic collapse occurs in the populations of herbivorous fish and coral, with the population of algae experiencing a ``bloom". Alternatively, we may observe ``tracking" of a sustainable coexistence state between the three populations in the presence of a folded focus.

math.DS

Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding

While 3DGS has emerged as a high-fidelity scene representation, encoding rich, general-purpose features directly from its primitives remains under-explored. We address this gap by introducing Chorus, a multi-teacher pretraining framework that learns a holistic feed-forward 3D Gaussian Splatting (3DGS) scene encoder by distilling complementary signals from 2D foundation models. Chorus employs a shared 3D encoder and teacher-specific projectors to learn from language-aligned, generalist, and object-aware teachers, encouraging a shared embedding space that captures signals from high-level semantics to fine-grained structure. We evaluate Chorus on a wide range of tasks: open-vocabulary semantic and instance segmentation, linear and decoder probing, data-efficient supervision, as well as LLM-based Q&A. Besides 3DGS, we also test Chorus on several benchmarks that only support point clouds by pretraining a variant using only Gaussian centers, colors, and estimated normals. Surprisingly, this encoder shows strong transfer and outperforms the point-cloud baseline while using 39.9 times fewer training scenes. Finally, we propose a render-and-distill adaptation that facilitates out-of-domain finetuning.

cs.CV

Effects of Model Reduction on Coherence and Information Transfer in Stochastic Biochemical Systems

Simplified stochastic models are widely used in the study of frequency-resolved noise propagation in biochemical reaction networks, a common measure being the coherence between random fluctuations in molecule number trajectories. Such models have also found widespread application in the quantification of how information is transmitted in reaction networks via the mutual information (MI) rate. A common assumption is that, under timescale separation, estimates for the coherence and MI rate obtained from simplified (reduced) models closely approximate those in the underlying full models. Here, we challenge that assumption by showing that, while reduced models can faithfully reproduce low-order statistics of molecular counts, they frequently incur substantial discrepancies in the coherence spectrum, especially at intermediate and high frequencies. These errors, in turn, lead to significant inaccuracies in the resulting estimates for the MI rates. We show that the observed discrepancies are due to the interplay between the structure of the underlying reaction networks, the specific model reduction method that is applied, and the asymptotic limits relating the full and the reduced models. We illustrate our results in canonical models of enzyme catalysis and gene expression, highlighting practical implications for quantifying information flow in cells.

q-bio.MN

SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting

3D Gaussian Splatting (3DGS) serves as a highly performant and efficient encoding of scene geometry, appearance, and semantics. Moreover, grounding language in 3D scenes has proven to be an effective strategy for 3D scene understanding. Current Language Gaussian Splatting line of work fall into three main groups: (i) per-scene optimization-based, (ii) per-scene optimization-free, and (iii) generalizable approach. However, most of them are evaluated only on rendered 2D views of a handful of scenes and viewpoints close to the training views, limiting ability and insight into holistic 3D understanding. To address this gap, we propose the first large-scale benchmark that systematically assesses these three groups of methods directly in 3D space, evaluating on 1060 scenes across three indoor datasets and one outdoor dataset. Benchmark results demonstrate a clear advantage of the generalizable paradigm, particularly in relaxing the scene-specific limitation, enabling fast feed-forward inference on novel scenes, and achieving superior segmentation performance. We further introduce GaussianWorld-49K a carefully curated 3DGS dataset comprising around 49K diverse indoor and outdoor scenes obtained from multiple sources, with which we demonstrate the generalizable approach could harness strong data priors. Our codes, benchmark, and datasets are released at https://scenesplatpp.gaussianworld.ai/.

cs.CV

SceneSplat: Gaussian Splatting-based Scene Understanding with Vision-Language Pretraining

Recognizing arbitrary or previously unseen categories is essential for comprehensive real-world 3D scene understanding. Currently, all existing methods rely on 2D or textual modalities during training or together at inference. This highlights the clear absence of a model capable of processing 3D data alone for learning semantics end-to-end, along with the necessary data to train such a model. Meanwhile, 3D Gaussian Splatting (3DGS) has emerged as the de facto standard for 3D scene representation across various vision tasks. However, effectively integrating semantic reasoning into 3DGS in a generalizable manner remains an open challenge. To address these limitations, we introduce SceneSplat, to our knowledge the first large-scale 3D indoor scene understanding approach that operates natively on 3DGS. Furthermore, we propose a self-supervised learning scheme that unlocks rich 3D feature learning from unlabeled scenes. To power the proposed methods, we introduce SceneSplat-7K, the first large-scale 3DGS dataset for indoor scenes, comprising 7916 scenes derived from seven established datasets, such as ScanNet and Matterport3D. Generating SceneSplat-7K required computational resources equivalent to 150 GPU days on an L4 GPU, enabling standardized benchmarking for 3DGS-based reasoning for indoor scenes. Our exhaustive experiments on SceneSplat-7K demonstrate the significant benefit of the proposed method over the established baselines.

cs.CV

The Burgers-FKPP advection-reaction-diffusion equation with cut-off

We investigate the effect of a Heaviside cut-off on the front propagation dynamics of the so-called Burgers-FisherKolmogoroff-Petrowskii-Piscounov (Burgers-FKPP) advection-reaction-diffusion equation. We prove the existence and uniqueness of a travelling front solution in the presence of a cut-off in the reaction kinetics and the advection term, and we derive the leading-order asymptotics for the speed of propagation of the front in dependence on the advection strength and the cut-off parameter. Our analysis relies on geometric techniques from dynamical systems theory and specifically, on geometric desingularisation, which also known as blow-up.

math.DS

Leveraging Driver Field-of-View for Multimodal Ego-Trajectory Prediction

Understanding drivers' decision-making is crucial for road safety. Although predicting the ego-vehicle's path is valuable for driver-assistance systems, existing methods mainly focus on external factors like other vehicles' motions, often neglecting the driver's attention and intent. To address this gap, we infer the ego-trajectory by integrating the driver's gaze and the surrounding scene. We introduce RouteFormer, a novel multimodal ego-trajectory prediction network combining GPS data, environmental context, and the driver's field-of-view, comprising first-person video and gaze fixations. We also present the Path Complexity Index (PCI), a new metric for trajectory complexity that enables a more nuanced evaluation of challenging scenarios. To tackle data scarcity and enhance diversity, we introduce GEM, a comprehensive dataset of urban driving scenarios enriched with synchronized driver field-of-view and gaze data. Extensive evaluations on GEM and DR(eye)VE demonstrate that RouteFormer significantly outperforms state-of-the-art methods, achieving notable improvements in prediction accuracy across diverse conditions. Ablation studies reveal that incorporating driver field-of-view data yields significantly better average displacement error, especially in challenging scenarios with high PCI scores, underscoring the importance of modeling driver attention. All data and code are available at https://meakbiyik.github.io/routeformer.

cs.CV

Model-aware 3D Eye Gaze from Weak and Few-shot Supervisions

The task of predicting 3D eye gaze from eye images can be performed either by (a) end-to-end learning for image-to-gaze mapping or by (b) fitting a 3D eye model onto images. The former case requires 3D gaze labels, while the latter requires eye semantics or landmarks to facilitate the model fitting. Although obtaining eye semantics and landmarks is relatively easy, fitting an accurate 3D eye model on them remains to be very challenging due to its ill-posed nature in general. On the other hand, obtaining large-scale 3D gaze data is cumbersome due to the required hardware setups and computational demands. In this work, we propose to predict 3D eye gaze from weak supervision of eye semantic segmentation masks and direct supervision of a few 3D gaze vectors. The proposed method combines the best of both worlds by leveraging large amounts of weak annotations--which are easy to obtain, and only a few 3D gaze vectors--which alleviate the difficulty of fitting 3D eye models on the semantic segmentation of eye images. Thus, the eye gaze vectors, used in the model fitting, are directly supervised using the few-shot gaze labels. Additionally, we propose a transformer-based network architecture, that serves as a solid baseline for our improvements. Our experiments in diverse settings illustrate the significant benefits of the proposed method, achieving about 5 degrees lower angular gaze error over the baseline, when only 0.05% 3D annotations of the training images are used. The source code is available at https://github.com/dimitris-christodoulou57/Model-aware_3D_Eye_Gaze.

cs.CV

Surface Normal Clustering for Implicit Representation of Manhattan Scenes

Novel view synthesis and 3D modeling using implicit neural field representation are shown to be very effective for calibrated multi-view cameras. Such representations are known to benefit from additional geometric and semantic supervision. Most existing methods that exploit additional supervision require dense pixel-wise labels or localized scene priors. These methods cannot benefit from high-level vague scene priors provided in terms of scenes' descriptions. In this work, we aim to leverage the geometric prior of Manhattan scenes to improve the implicit neural radiance field representations. More precisely, we assume that only the knowledge of the indoor scene (under investigation) being Manhattan is known -- with no additional information whatsoever -- with an unknown Manhattan coordinate frame. Such high-level prior is used to self-supervise the surface normals derived explicitly in the implicit neural fields. Our modeling allows us to cluster the derived normals and exploit their orthogonality constraints for self-supervision. Our exhaustive experiments on datasets of diverse indoor scenes demonstrate the significant benefit of the proposed method over the established baselines. The source code is available at https://github.com/nikola3794/normal-clustering-nerf.

cs.CV

Improving the Behaviour of Vision Transformers with Token-consistent Stochastic Layers

We introduce token-consistent stochastic layers in vision transformers, without causing any severe drop in performance. The added stochasticity improves network calibration, robustness and strengthens privacy. We use linear layers with token-consistent stochastic parameters inside the multilayer perceptron blocks, without altering the architecture of the transformer. The stochastic parameters are sampled from the uniform distribution, both during training and inference. The applied linear operations preserve the topological structure, formed by the set of tokens passing through the shared multilayer perceptron. This operation encourages the learning of the recognition task to rely on the topological structures of the tokens, instead of their values, which in turn offers the desired robustness and privacy of the visual features. The effectiveness of the token-consistent stochasticity is demonstrated on three different applications, namely, network calibration, adversarial robustness, and feature privacy, by boosting the performance of the respective established baselines.

cs.CV

Spatially Multi-conditional Image Generation

In most scenarios, conditional image generation can be thought of as an inversion of the image understanding process. Since generic image understanding involves solving multiple tasks, it is natural to aim at generating images via multi-conditioning. However, multi-conditional image generation is a very challenging problem due to the heterogeneity and the sparsity of the (in practice) available conditioning labels. In this work, we propose a novel neural architecture to address the problem of heterogeneity and sparsity of the spatially multi-conditional labels. Our choice of spatial conditioning, such as by semantics and depth, is driven by the promise it holds for better control of the image generation process. The proposed method uses a transformer-like architecture operating pixel-wise, which receives the available labels as input tokens to merge them in a learned homogeneous space of labels. The merged labels are then used for image generation via conditional generative adversarial training. In this process, the sparsity of the labels is handled by simply dropping the input tokens corresponding to the missing labels at the desired locations, thanks to the proposed pixel-wise operating architecture. Our experiments on three benchmark datasets demonstrate the clear superiority of our method over the state-of-the-art and compared baselines. The source code will be made publicly available.

cs.CV

Gradient Obfuscation Checklist Test Gives a False Sense of Security

One popular group of defense techniques against adversarial attacks is based on injecting stochastic noise into the network. The main source of robustness of such stochastic defenses however is often due to the obfuscation of the gradients, offering a false sense of security. Since most of the popular adversarial attacks are optimization-based, obfuscated gradients reduce their attacking ability, while the model is still susceptible to stronger or specifically tailored adversarial attacks. Recently, five characteristics have been identified, which are commonly observed when the improvement in robustness is mainly caused by gradient obfuscation. It has since become a trend to use these five characteristics as a sufficient test, to determine whether or not gradient obfuscation is the main source of robustness. However, these characteristics do not perfectly characterize all existing cases of gradient obfuscation, and therefore can not serve as a basis for a conclusive test. In this work, we present a counterexample, showing this test is not sufficient for concluding that gradient obfuscation is not the main cause of improvements in robustness.

cs.CV

CompositeTasking: Understanding Images by Spatial Composition of Tasks

We define the concept of CompositeTasking as the fusion of multiple, spatially distributed tasks, for various aspects of image understanding. Learning to perform spatially distributed tasks is motivated by the frequent availability of only sparse labels across tasks, and the desire for a compact multi-tasking network. To facilitate CompositeTasking, we introduce a novel task conditioning model -- a single encoder-decoder network that performs multiple, spatially varying tasks at once. The proposed network takes an image and a set of pixel-wise dense task requests as inputs, and performs the requested prediction task for each pixel. Moreover, we also learn the composition of tasks that needs to be performed according to some CompositeTasking rules, which includes the decision of where to apply which task. It not only offers us a compact network for multi-tasking, but also allows for task-editing. Another strength of the proposed method is demonstrated by only having to supply sparse supervision per task. The obtained results are on par with our baselines that use dense supervision and a multi-headed multi-tasking design. The source code will be made publicly available at www.github.com/nikola3794/composite-tasking.

cs.CV

Rethinking Global Context in Crowd Counting

This paper investigates the role of global context for crowd counting. Specifically, a pure transformer is used to extract features with global information from overlapping image patches. Inspired by classification, we add a context token to the input sequence, to facilitate information exchange with tokens corresponding to image patches throughout transformer layers. Due to the fact that transformers do not explicitly model the tried-and-true channel-wise interactions, we propose a token-attention module (TAM) to recalibrate encoded features through channel-wise attention informed by the context token. Beyond that, it is adopted to predict the total person count of the image through regression-token module (RTM). Extensive experiments on various datasets, including ShanghaiTech, UCF-QNRF, JHU-CROWD++ and NWPU, demonstrate that the proposed context extraction techniques can significantly improve the performance over the baselines.

cs.CV

Jump-induced mixed-mode oscillations through piecewise-affine maps

Mixed-mode oscillations (MMOs) are complex oscillatory patterns in which large-amplitude relaxation oscillations (LAOs) alternate with small-amplitude oscillations (SAOs). MMOs are found in singularly perturbed systems of ordinary differential equations of slow-fast type, and are typically related to the presence of so-called folded singularities and the corresponding canard trajectories in such systems. Here, we introduce a canonical family of three-dimensional slow-fast systems that exhibit MMOs which are induced by relaxation-type dynamics, and which are hence based on a "jump mechanism", rather than on a more standard canard mechanism. In particular, we establish a correspondence between that family and a class of associated one-dimensional piecewise affine maps (PAMs) which exhibit MMOs with the same signature. Finally, we give a preliminary classification of admissible mixed-mode signatures, and we illustrate our findings with numerical examples.

math.DS

Singular perturbation analysis of a regularized MEMS model

Micro-Electro Mechanical Systems (MEMS) are defined as very small structures that combine electrical and mechanical components on a common substrate. Here, the electrostatic-elastic case is considered, where an elastic membrane is allowed to deflect above a ground plate under the action of an electric potential, whose strength is proportional to a parameter $\lambda$. Such devices are commonly described by a parabolic partial differential equation that contains a singular nonlinear source term. The singularity in that term corresponds to the so-called "touchdown" phenomenon, where the membrane establishes contact with the ground plate. Touchdown is known to imply the non-existence of steady state solutions and blow-up of solutions in finite time. We study a recently proposed extension of that canonical model, where such singularities are avoided due to the introduction of a regularizing term involving a small "regularization" parameter $\varepsilon$. Methods from dynamical systems and geometric singular perturbation theory, in particular the desingularization technique known as "blow-up", allow for a precise description of steady-state solutions of the regularized model, as well as for a detailed resolution of the resulting bifurcation diagram. The interplay between the two main model parameters $\varepsilon$ and $\lambda$ is emphasized; in particular, the focus is on the singular limit as both parameters tend to zero.

math.DS