SearcharxivSearch

arXiv subjects

Dun Liang

Publications and source records attributed to Dun Liang.

12 recordsLinked to original sources

SciDocBench: A Workflow-Centered Benchmark and Data Pipeline for Scientific Document Understanding

Scientific papers require models to reason jointly over text, equations, figures, tables, code, and datasets while preserving the provenance of supporting evidence. Existing benchmarks typically evaluate these capabilities in isolation, leaving unclear whether multimodal models can support realistic scientific-reading workflows. We introduce SciDocBench, a workflow-centered benchmark for scientific document understanding. It contains 124 expert-authored and difficulty-screened questions organized into seven research-assistant capability groups and 19 subtasks across five scientific domains. Each question is instantiated under four matched conditions combining English or Chinese questions with all-images-first or interleaved document representations, yielding 496 evaluation instances for controlled analysis. The strongest evaluated system achieves only 62.6/100, with pronounced weaknesses in document perception, evidence grounding, verification, and cross-document reasoning. To translate these diagnostics into scalable training signals, we introduce SciDocIR, a typed evidence-graph representation that preserves scientific document objects, layout and cross-reference relations, and provenance. Building on SciDocIR, we construct SciDocDataset, comprising approximately 15K supervised fine-tuning samples and 8K reinforcement-learning samples across 14 verifiable subtasks. Together, SciDocBench, SciDocIR, and SciDocDataset form an evaluation-to-training framework for diagnosing and improving scientific-document assistants. The project page is available at https://github.com/InternLM/SciDocBench.

cs.AI

Limit Values of Character Sums in Frobenius Formula of Three Permutations

We study the asymptotic behaviour of the character summation part of the Frobenius formula for three conjugacy classes of the symmetric groups. When two conjugacy classes contain fixed points of order $\sim H_i\sqrt{n}$ for $i=1,2$ and all other cycles are long, the summation converges uniformly to $2e^{-H_1H_2}$; the exponential factor coincides with the non-collision probability of the birthday paradox. When short cycles are rare in all three classes, character sum tends to $2$.

math.GM

Are we ready for a new paradigm shift? A Survey on Visual Deep MLP

Recently, the proposed deep MLP models have stirred up a lot of interest in the vision community. Historically, the availability of larger datasets combined with increased computing capacity leads to paradigm shifts. This review paper provides detailed discussions on whether MLP can be a new paradigm for computer vision. We compare the intrinsic connections and differences between convolution, self-attention mechanism, and Token-mixing MLP in detail. Advantages and limitations of Token-mixing MLP are provided, followed by careful analysis of recent MLP-like variants, from module design to network architecture, and their applications. In the GPU era, the locally and globally weighted summations are the current mainstreams, represented by the convolution and self-attention mechanism, as well as MLP. We suggest the further development of paradigm to be considered alongside the next-generation computing devices.

cs.CV

Can Attention Enable MLPs To Catch Up With CNNs?

In the first week of May, 2021, researchers from four different institutions: Google, Tsinghua University, Oxford University and Facebook, shared their latest work [16, 7, 12, 17] on arXiv.org almost at the same time, each proposing new learning architectures, consisting mainly of linear layers, claiming them to be comparable, or even superior to convolutional-based models. This sparked immediate discussion and debate in both academic and industrial communities as to whether MLPs are sufficient, many thinking that learning architectures are returning to MLPs. Is this true? In this perspective, we give a brief history of learning architectures, including multilayer perceptrons (MLPs), convolutional neural networks (CNNs) and transformers. We then examine what the four newly proposed architectures have in common. Finally, we give our views on challenges and directions for new learning architectures, hoping to inspire future research.

cs.CV

Recursive-NeRF: An Efficient and Dynamically Growing NeRF

View synthesis methods using implicit continuous shape representations learned from a set of images, such as the Neural Radiance Field (NeRF) method, have gained increasing attention due to their high quality imagery and scalability to high resolution. However, the heavy computation required by its volumetric approach prevents NeRF from being useful in practice; minutes are taken to render a single image of a few megapixels. Now, an image of a scene can be rendered in a level-of-detail manner, so we posit that a complicated region of the scene should be represented by a large neural network while a small neural network is capable of encoding a simple region, enabling a balance between efficiency and quality. Recursive-NeRF is our embodiment of this idea, providing an efficient and adaptive rendering and training approach for NeRF. The core of Recursive-NeRF learns uncertainties for query coordinates, representing the quality of the predicted color and volumetric intensity at each level. Only query coordinates with high uncertainties are forwarded to the next level to a bigger neural network with a more powerful representational capability. The final rendered image is a composition of results from neural networks of all levels. Our evaluation on three public datasets shows that Recursive-NeRF is more efficient than NeRF while providing state-of-the-art quality. The code will be available at https://github.com/Gword/Recursive-NeRF.

cs.CV

Invariants, Bitangents and Matrix Representations of Plane Quartics with 3-Cyclic Automorphisms

In this work we compute the Dixmier invariants and bitangents of the plane quartics with 3,6 or 9-cyclic automorphisms, we find that a quartic curve with 6-cyclic automorphism will have 3 horizontal bitangents which form an asysgetic triple. We also discuss the linear matrix representation problem of such curves, and find a degree 6 equation of 1 variable which solves the symbolic solution of the linear matrix representation problem for the curve with 6-cyclic automorphism.

math.AG

What and Where: A Context-based Recommendation System for Object Insertion

In this work, we propose a novel topic consisting of two dual tasks: 1) given a scene, recommend objects to insert, 2) given an object category, retrieve suitable background scenes. A bounding box for the inserted object is predicted in both tasks, which helps downstream applications such as semi-automated advertising and video composition. The major challenge lies in the fact that the target object is neither present nor localized at test time, whereas available datasets only provide scenes with existing objects. To tackle this problem, we build an unsupervised algorithm based on object-level contexts, which explicitly models the joint probability distribution of object categories and bounding boxes with a Gaussian mixture model. Experiments on our newly annotated test set demonstrate that our system outperforms existing baselines on all subtasks, and do so under a unified framework. Our contribution promises future extensions and applications.

cs.CV

Algorithms And Programming On The Minimal Combinations Of Weights Of Projective Hypersurfaces

This paper designs an alogrithm to compute the minimal combinations of finite sets in Euclidean spaces, and applys the algorithm of study the moment maps and geometric invariant stability of hypersurfaces. The classical example of cubic curves is repeated by the algorithm. Furhtermore the alogrithm works for cubic surfaces. For given affinely indepdent subsets of monomials, the algorithm can output the unique unstable points of the Morse strata if it exists. Also there is a discussion on the affinely dependent sets of monomials.

math.AG

LineNet: a Zoomable CNN for Crowdsourced High Definition Maps Modeling in Urban Environments

High Definition (HD) maps play an important role in modern traffic scenes. However, the development of HD maps coverage grows slowly because of the cost limitation. To efficiently model HD maps, we proposed a convolutional neural network with a novel prediction layer and a zoom module, called LineNet. It is designed for state-of-the-art lane detection in an unordered crowdsourced image dataset. And we introduced TTLane, a dataset for efficient lane detection in urban road modeling applications. Combining LineNet and TTLane, we proposed a pipeline to model HD maps with crowdsourced data for the first time. And the maps can be constructed precisely even with inaccurate crowdsourced data.

cs.CV

Computing Moment Maps of Hypersurfaces using MAXIMA

We use Maxima to compute the moment matrices of hypersurfaces. After that, we compute the Hilbert-Mumford numerical criterion for plane cubics and plane quartics, and give the stability of these curves.

math.AG

Genus 3 curves whose Jacobians have endomorphisms by $Q (\zeta _7 +\bar{\zeta}_7 )$, II

In this work we consider constructions of genus three curves $X$ such that $\mathrm{End}(\mathrm{Jac} (X))\otimes Q$ contains the totally real cubic number field $Q(\zeta _7 +\bar{\zeta}_7 )$. We construct explicit three-dimensional families whose generic member is a nonhyperelliptic genus 3 curve with this property. The case when $X$ is hyperelliptic was studied in a previous work by Hoffman and Wang and some nonhyperelliptic curves were constructed in a previous paper by Hoffman, Z. Liang. Sakai and Wang.

math.AG