SearcharxivSearch

arXiv subjects

Xintian Zhang

Publications and source records attributed to Xintian Zhang.

6 recordsLinked to original sources

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Pretraining scaling laws reveal that model capability improves predictably with data and compute. But learning from real world environments after deployment remains far less understood. Analyzing roughly 38,000 hours of agent interaction with the environment across 134 real world tasks, we find, to the best of our knowledge, the first evidence that overall performance during environment learning follows a log-sigmoid scaling law with remarkably high precision, reaching R^2 = 0.998. Across model generations, we also find that agent learning speed roughly doubles every three months. This discovery stems from EdgeBench, a suite of 134 real world tasks with ultra-long horizons, spanning scientific discovery, software engineering, combinatorial optimization, professional knowledge work, formal mathematics, and interactive games. Each task sustains at least 12 hours of continuous agent operation under rich, multilevel feedback, and is built through substantial expert effort. We publicly release 51 tasks and our full evaluation framework to accelerate the study of how agents learn from real world experience.

cs.CL

ViewSAM: Learning View-aware Cross-modal Semantics for Weakly Supervised Cross-view Referring Multi-Object Tracking

Cross-view Referring Multi-Object Tracking (CRMOT) aims to track multiple objects specified by natural language across multiple camera views, with globally consistent identities. Despite recent progress, existing methods rely heavily on costly frame-level spatial annotations and cross-view identity supervision. To reduce such reliance, we explore CRMOT under weak supervision by leveraging the capabilities of foundation models. However, our empirical study shows that directly applying foundation models such as SAM2 and SAM3, even with task-specific modifications, fails to accurately understand referring expressions and maintain consistent identities across views. Yet, they remain effective at producing reliable object tracklets that can serve as pseudo supervision. We therefore repurpose foundation models as pseudo-label generators and propose a two-stage framework for weakly supervised CRMOT, using only object category labels as coarse-grained supervision. In the first stage, we design an Affinity-guided Cross-view Re-prompting strategy to refine and associate SAM3-generated tracklets across cameras, producing reliable cross-view pseudo labels for subsequent training. In the second stage, we introduce ViewSAM, a CRMOT model built upon SAM2 that explicitly models view-aware cross-modal semantics. By formulating view-induced variations as learnable conditions, ViewSAM bridges the gap between view-variant visual observations and view-invariant textual expressions, enabling robust cross-view referring tracking with only approximately 10% additional parameters. Extensive experiments demonstrate that ViewSAM achieves SOTA performance under weak supervision and remains competitive with fully supervised methods.

cs.CV

On the Fourier transform of random Bernoulli convolutions

We investigate random Bernoulli convolutions, namely, probability measures given by the infinite convolution \[ μ_ω= \mathop{\circledast}_{k=1}^{\infty} \left( \frac{δ_0 + δ_{λ_1 λ_2 \ldots λ_{k-1} λ_k}}{2} \right), \] where $ω=(λ_k)$ is a sequence of i.i.d. random variables each following the uniform distribution on some fixed interval. We study the regularity of these measures and prove that when $\exp\mathbb{E}\left( \log λ_1\right)>\frac{2}π, $ the Fourier transform $\widehatμ_ω$ is an $L^{1}$ function almost surely. This in turn implies that the corresponding random self-similar set supporting $μ_ω$ has non-empty interior almost surely. This improves upon a previous bound due to Peres, Simon and Solomyak. Furthermore, under no assumptions on the value of $\exp \mathbb{E}(\log λ_1), $ we prove that $\widehat μ_ω$ will decay to zero at a polynomial rate almost surely.

math.DS

Transfer Learning for Assessing Heavy Metal Pollution in Seaports Sediments

Detecting heavy metal pollution in soils and seaports is vital for regional environmental monitoring. The Pollution Load Index (PLI), an international standard, is commonly used to assess heavy metal containment. However, the conventional PLI assessment involves laborious procedures and data analysis of sediment samples. To address this challenge, we propose a deep-learning-based model that simplifies the heavy metal assessment process. Our model tackles the issue of data scarcity in the water-sediment domain, which is traditionally plagued by challenges in data collection and varying standards across nations. By leveraging transfer learning, we develop an accurate quantitative assessment method for predicting PLI. Our approach allows the transfer of learned features across domains with different sets of features. We evaluate our model using data from six major ports in New South Wales, Australia: Port Yamba, Port Newcastle, Port Jackson, Port Botany, Port Kembla, and Port Eden. The results demonstrate significantly lower Mean Absolute Error (MAE) and Mean Absolute Percentage Error (MAPE) of approximately 0.5 and 0.03, respectively, compared to other models. Our model performance is up to 2 orders of magnitude than other baseline models. Our proposed model offers an innovative, accessible, and cost-effective approach to predicting water quality, benefiting marine life conservation, aquaculture, and industrial pollution monitoring.

cs.LG

Hausdorff dimension of the set of eventually always hitting points on a self-conformal set

Recurrence problems are fundamental in dynamics, and for example, sizes of the set of points recurring infinitely often to a target have been studied extensively in many contexts. For example, the problem of finding the dimension for shrinking target set in an iterated function system is an active research area. In the current work, we consider a set with a finer recurrence quality, the eventually always hitting set. In a sense, the points in the intersection of an eventually always hitting set and a shrinking target set not only return infinitely often but also at a bounded rate. We study this set in the context of self-conformal iterated function systems, and compute upper and lower bounds for its Hausdorff dimension. Additionally, as an intermediate theorem, we obtain a Hausdorff dimension result for the intersection of eventually always hitting and shrinking target sets.

math.DS

The dimension of the set of $ψ$-badly approximable points in all ambient dimensions; on a question of Beresnevich and Velani

Let $ψ:\mathbb{N} \to [0,\infty)$, $ψ(q)=q^{-(1+τ)}$ and let $ψ$-badly approximable points be those vectors in $\mathbb{R}^{d}$ that are $ψ$-well approximable, but not $cψ$-well approximable for arbitrarily small constants $c>0$. We establish that the $ψ$-badly approximable points have the Hausdorff dimension of the $ψ$-well approximable points, the dimension taking the value $(d+1)/(τ+1)$ familiar from theorems of Besicovitch and Jarník. The method of proof is an entirely new take on the Mass Transference Principle by Beresnevich and Velani (Annals, 2006); namely, we use the colloquially named `delayed pruning' to construct a sufficiently large $\liminf$ set and combine this with ideas inspired by the proof of the Mass Transference Principle to find a large $\limsup$ subset of the $\liminf$ set. Our results are a generalisation of some $1$-dimensional results due to Bugeaud and Moreira (Acta Arith, 2011), but our method of proof is nothing alike.

math.NT