SearcharxivSearch

arXiv subjects

Jinha Kim

Publications and source records attributed to Jinha Kim.

At least 19 recordsLinked to original sources

Retrofitting Linear Attention into Diffusion Language Models

Diffusion language models (dLLMs) offer a promising alternative to autoregressive models by accelerating inference through parallel decoding. Recent dLLMs commonly use blockwise semi-autoregressive decoding, generating blocks autoregressively while denoising tokens within each active block in parallel. However, despite KV caching, each denoising step still attends to all previous blocks, repeatedly incurring prefix-attention cost. Motivated by this bottleneck, we ask whether dLLM inference can be further accelerated by linearizing attention over previous blocks. We introduce block-hybrid attention, which retains exact softmax attention within the active denoising block while applying linear attention over previous blocks. We show that this hybrid attention can be retrofitted into a pretrained dLLM with minimal post-training: LLaDA-Hybrid replaces 6 of the 20 attention layers in LLaDA~2.1, a 16B open-source dLLM, largely following LoLCAT (Zhang et al, 2024). The conversion takes only approximately 60 hours while preserving benchmark performance: 72.0% vs. 75.6% on HumanEval, 63.0% vs. 57.7% on MBPP+, and 86.7% vs. 88.3% on CMATH. With a Triton implementation, LLaDA-Hybrid achieves up to $1.7\times$ higher decoding throughput and supports more concurrent requests before exhausting memory, showing that pretrained dLLMs can be efficiently linearized for faster inference. Our code is available at: https://github.com/Diuven/LLaDA-Hybrid.

cs.LG

Domination numbers and homotopy in certain ternary graphs

A ternary graph is a graph with no induced cycles of length $0$ modulo $3$. It was recently shown that, if the independence complex of a ternary graph is not contractible, then it is homotopy equivalent to a sphere. When a ternary graph also does not contain induced cycles of length $1$ modulo $3$, we prove that the dimension of the sphere is equal to the dimension of a minimum maximal simplex of the independence complex, or equivalently, to the value obtained by subtracting $1$ from the independent domination number of the graph. The same statement holds if we replace the independent domination number with the domination number. We also give a hypergraph analogue of the statement above.

math.CO

WACA-UNet: Weakness-Aware Channel Attention for Static IR Drop Prediction in Integrated Circuit Design

Accurate spatial prediction of power integrity issues, such as IR drop, is critical for reliable VLSI design. However, traditional simulation-based solvers are computationally expensive and difficult to scale. We address this challenge by reformulating IR drop estimation as a pixel-wise regression task on heterogeneous multi-channel physical maps derived from circuit layouts. Prior learning-based methods treat all input layers (e.g., metal, via, and current maps) equally, ignoring their varying importance to prediction accuracy. To tackle this, we propose a novel Weakness-Aware Channel Attention (WACA) mechanism, which recursively enhances weak feature channels while suppressing over-dominant ones through a two-stage gating strategy. Integrated into a ConvNeXtV2-based attention U-Net, our approach enables adaptive and balanced feature representation. On the public ICCAD-2023 benchmark, our method outperforms the ICCAD-2023 contest winner by reducing mean absolute error by 61.1% and improving F1-score by 71.0%. These results demonstrate that channel-wise heterogeneity is a key inductive bias in physical layout analysis for VLSI.

cs.LG

Colorful fractional Helly theorem via weak saturation

Two celebrated extensions of the classical Helly's theorem are the fractional Helly theorem and the colorful Helly theorem. Bulavka, Goodarzi, and Tancer recently established the optimal bound for the unified generalization of the fractional and the colorful Helly theorems using a colored extension of the exterior algebra. In this paper, we combinatorially reduce both the fractional Helly theorem and its colorful version to a classical problem in extremal combinatorics known as {weak saturation}. No such results connecting the fractional Helly theorem and weak saturation are known in the long history of literature. These reductions, along with basic linear algebraic arguments for the reduced weak saturation problems, let us give new short proofs of the optimal bounds for both the fractional Helly theorem and its colorful version without using exterior algebra.

math.CO

Topology of independence complexes and cycle structure of hypergraphs

Recently, Zhang and Wu proved a conjecture of Kalai and Meshulam, showing that for every graph $G$ without induced cycles of length divisible by $3$, the sum of all reduced Betti numbers of its independence complex $I(G)$ is at most $1$. We extend this result to the hypergraph setting. Namely, we show that the same conclusion holds for any hypergraph $H$ that does not contain a Berge cycle of length divisible by $3$. This establishes a broader connection between forbidden cycle structures and the topological simplicity of independence complexes. As a key tool, we introduce a hypergraph analogue of Barmak's star cluster theorem for graphs. This new theorem implies, in particular, that if a hypergraph $H$ has a vertex $v$ that is not isolated and is not contained in an induced Berge cycle of length $3$, then there exists a hypergraph $H'$ with fewer vertices than $H$ such that the independence complex of $H$ is homotopy equivalent to the suspension of the independence complex of $H'$.

math.CO

Swish-T : Enhancing Swish Activation with Tanh Bias for Improved Neural Network Performance

We propose the Swish-T family, an enhancement of the existing non-monotonic activation function Swish. Swish-T is defined by adding a Tanh bias to the original Swish function. This modification creates a family of Swish-T variants, each designed to excel in different tasks, showcasing specific advantages depending on the application context. The Tanh bias allows for broader acceptance of negative values during initial training stages, offering a smoother non-monotonic curve than the original Swish. We ultimately propose the Swish-T$_{\textbf{C}}$ function, while Swish-T and Swish-T$_{\textbf{B}}$, byproducts of Swish-T$_{\textbf{C}}$, also demonstrate satisfactory performance. Furthermore, our ablation study shows that using Swish-T$_{\textbf{C}}$ as a non-parametric function can still achieve high performance. The superiority of the Swish-T family has been empirically demonstrated across various models and benchmark datasets, including MNIST, Fashion MNIST, SVHN, CIFAR-10, and CIFAR-100. The code is publicly available at https://github.com/ictseoyoungmin/Swish-T-pytorch.

cs.LG

Can Contrastive Learning Refine Embeddings

Recent advancements in contrastive learning have revolutionized self-supervised representation learning and achieved state-of-the-art performance on benchmark tasks. While most existing methods focus on applying contrastive learning to input data modalities such as images, natural language sentences, or networks, they overlook the potential of utilizing outputs from previously trained encoders. In this paper, we introduce SIMSKIP, a novel contrastive learning framework that specifically refines input embeddings for downstream tasks. Unlike traditional unsupervised learning approaches, SIMSKIP takes advantage of the output embeddings of encoder models as its input. Through theoretical analysis, we provide evidence that applying SIMSKIP does not result in larger upper bounds on downstream task errors than those of the original embeddings, which serve as SIMSKIP's input. Experimental results on various open datasets demonstrate that the embeddings produced by SIMSKIP improve performance on downstream tasks.

cs.LG

An Eisenbud-Goto type inequality for Stanley-Reisner ideals and simplicial complexes

The Leray number of an abstract simplicial complex is the minimal integer $d$ where its induced subcomplexes have trivial homology groups in dimension $d$ or greater. We give an upper bound on the Leray number of a complex in terms of how the facets are attached to each other. We also describe the structure of complexes for the equality of the bound that we found. Through the Stanley-Reisner correspondence, our results give an Eisenbud-Goto type inequality for any square-free monomial ideals. This generalizes Terai's result.

math.AC

Transversal numbers of stacked spheres

A stacked $d$-sphere $S$ is the boundary complex of a stacked $(d+1)$-ball, which is obtained by taking cone over a free $d$-face repeatedly from a $(d+1)$-simplex. A stacked sphere $S$ is called linear if every cone is taken over a face added in the previous step. In this paper, we study the transversal number of facets of stacked $d$-spheres, denoted by $τ(S)$, which is the minimum number of vertices intersecting with all facets. Briggs, Dobbins and Lee showed that the transversal ratio of a stacked $d$-sphere is bounded above by $\frac{2}{d+2}+o(1)$ and can be as large as $\frac{2}{d+3}$. We improve the lower bound by constructing linear stacked $d$-spheres with transversal ratio $\frac{6}{3d+8}$ and general stacked $d$-spheres with transversal ratio $\frac{2d+3}{(d+2)^2}$. Finally, we show that $\frac{6}{3d+8}$ is optimal for linear stacked $2$-spheres, that is, the transversal ratio is at most $\frac{3}{7} + o(1)$ for linear stacked $2$-spheres.

math.CO

Strong Erdős-Hajnal properties in chordal graphs

A graph class $\mathcal{G}$ has the strong Erdős-Hajnal property (SEH-property) if there is a constant $c=c(\mathcal{G}) > 0$ such that for every member $G$ of $\mathcal{G}$, either $G$ or its complement has $K_{m, m}$ as a subgraph where $m \geq \left\lfloor c|V(G)|\right\rfloor$. We prove that the class of chordal graphs satisfy SEH-property with constant $c = 2/9$. On the other hand, a strengthening of SEH-property which we call the colorful Erdős-Hajnal property was discussed in geometric settings by Alon et al. (2005) and by Fox et al. (2012). Inspired by their results, we show that for every pair $F_1, F_2$ of subtree families of the same size in a tree $T$ with $k$ leaves, there exists subfamilies $F'_1 \subseteq F_1$ and $F'_2 \subseteq F_2$ of size $θ\left( \frac{\ln k}{k} \left| F_1 \right|\right)$ such that either every pair of representatives from distinct subfamilies intersect or every such pair do not intersect. Our results are asymptotically optimal.

math.CO

Independent domination of graphs with bounded maximum degree

An independent dominating set of a graph, also known as a maximal independent set, is a set $S$ of pairwise non-adjacent vertices such that every vertex not in $S$ is adjacent to some vertex in $S$. We prove that for $Δ=4$ or $Δ\ge 6$, every connected $n$-vertex graph of maximum degree at most $Δ$ has an independent dominating set of size at most $(1-\fracΔ{\lfloorΔ^2/4\rfloor+Δ})(n-1)+1$. In addition, we characterize all connected graphs having the equality and we show that other connected graphs have an independent dominating set of size at most $(1-\fracΔ{ \lfloorΔ^2/4\rfloor+Δ})n$.

math.CO

Reversed Image Signal Processing and RAW Reconstruction. AIM 2022 Challenge Report

Cameras capture sensor RAW images and transform them into pleasant RGB images, suitable for the human eyes, using their integrated Image Signal Processor (ISP). Numerous low-level vision tasks operate in the RAW domain (e.g. image denoising, white balance) due to its linear relationship with the scene irradiance, wide-range of information at 12bits, and sensor designs. Despite this, RAW image datasets are scarce and more expensive to collect than the already large and public RGB datasets. This paper introduces the AIM 2022 Challenge on Reversed Image Signal Processing and RAW Reconstruction. We aim to recover raw sensor images from the corresponding RGBs without metadata and, by doing this, "reverse" the ISP transformation. The proposed methods and benchmark establish the state-of-the-art for this low-level vision inverse problem, and generating realistic raw sensor readings can potentially benefit other tasks such as denoising and super-resolution.

eess.IV

Overexposure Mask Fusion: Generalizable Reverse ISP Multi-Step Refinement

With the advent of deep learning methods replacing the ISP in transforming sensor RAW readings into RGB images, numerous methodologies solidified into real-life applications. Equally potent is the task of inverting this process which will have applications in enhancing computational photography tasks that are conducted in the RAW domain, addressing lack of available RAW data while reaping from the benefits of performing tasks directly on sensor readings. This paper's proposed methodology is a state-of-the-art solution to the task of RAW reconstruction, and the multi-step refinement process integrating an overexposure mask is novel in three ways: instead of from RGB to bayer, the pipeline trains from RGB to demosaiced RAW allowing use of perceptual loss functions; the multi-step processes has greatly enhanced the performance of the baseline U-Net from start to end; the pipeline is a generalizable process of refinement that can enhance other high performance methodologies that support end-to-end learning.

cs.CV

Unified almost linear kernels for generalized covering and packing problems on nowhere dense classes

Let $\mathcal{F}$ be a family of graphs, and let $p,r$ be nonnegative integers. The \textsc{$(p,r,\mathcal{F})$-Covering} problem asks whether for a graph $G$ and an integer $k$, there exists a set $D$ of at most $k$ vertices in $G$ such that $G^p\setminus N_G^r[D]$ has no induced subgraph isomorphic to a graph in $\mathcal{F}$, where $G^p$ is the $p$-th power of $G$. The \textsc{$(p,r,\mathcal{F})$-Packing} problem asks whether for a graph $G$ and an integer $k$, $G^p$ has $k$ induced subgraphs $H_1,\ldots,H_k$ such that each $H_i$ is isomorphic to a graph in $\mathcal{F}$, and for distinct $i,j\in \{1, \ldots, k\}$, the distance between $V(H_i)$ and $V(H_j)$ in $G$ is larger than $r$. We show that for every fixed nonnegative integers $p,r$ and every fixed nonempty finite family $\mathcal{F}$ of connected graphs, the \textsc{$(p,r,\mathcal{F})$-Covering} problem with $p\leq2r+1$ and the \textsc{$(p,r,\mathcal{F})$-Packing} problem with $p\leq2\lfloor r/2\rfloor+1$ admit almost linear kernels on every nowhere dense class of graphs, and admit linear kernels on every class of graphs with bounded expansion, parameterized by the solution size $k$. We obtain the same kernels for their annotated variants. As corollaries, we prove that \textsc{Distance-$r$ Vertex Cover}, \textsc{Distance-$r$ Matching}, \textsc{$\mathcal{F}$-Free Vertex Deletion}, and \textsc{Induced-$\mathcal{F}$-Packing} for any fixed finite family $\mathcal{F}$ of connected graphs admit almost linear kernels on every nowhere dense class of graphs and linear kernels on every class of graphs with bounded expansion. Our results extend the results for \textsc{Distance-$r$ Dominating Set} by Drange et al. (STACS 2016) and Eickmeyer et al. (ICALP 2017), and the result for \textsc{Distance-$r$ Independent Set} by Pilipczuk and Siebertz (EJC 2021).

cs.DS

Well-mixing vertices and almost expanders

We study regular graphs in which the random walks starting from a positive fraction of vertices have small mixing time. We prove that any such graph is virtually an expander and has no small separator. This answers a question of Pak [SODA, 2002]. As a corollary, it shows that sparse (constant degree) regular graphs with many well-mixing vertices have a long cycle, improving a result of Pak. Furthermore, such cycle can be found in polynomial time. Secondly, we show that if the random walks from a positive fraction of vertices are well-mixing, then the random walks from almost all vertices are well-mixing (with a slightly worse mixing time).

math.CO

The homotopy type of the independence complex of graphs with no induced cycles of length divisible by $3$

We prove Engström's conjecture that the independence complex of graphs with no induced cycle of length divisible by $3$ is either contractible or homotopy equivalent to a sphere. Our result strengthens a result by Zhang and Wu, verifying a conjecture of Kalai and Meshulam which states that the total Betti number of the independence complex of such a graph is at most $1$. A weaker conjecture was proved earlier by Chudnovsky, Scott, Seymour, and Spirkl, who showed that in such a graph, the number of independent sets of even size minus the number of independent sets of odd size has values $0$, $1$, or $-1$.

math.CO

Cooperative conditions for the existence of rainbow matchings

Let $k>1$, and let $\mathcal{F}$ be a family of $2n+k-3$ non-empty sets of edges in a bipartite graph. If the union of every $k$ members of $\mathcal{F}$ contains a matching of size $n$, then there exists an $\mathcal{F}$-rainbow matching of size $n$. Replacing $2n+k-3$ by $2n+k-2$, the result is true also for $k=1$, and it can be proved (for all $k$) both topologically and by a relatively simple combinatorial argument. The main effort is in gaining the last $1$, which makes the result sharp.

math.CO

Fractional Helly theorem for Cartesian products of convex sets

Helly's theorem and its variants show that for a family of convex sets in Euclidean space, local intersection patterns influence global intersection patterns. A classical result of Eckhoff in 1988 provided an optimal fractional Helly theorem for axis-aligned boxes, which are Cartesian products of line segments. Answering a question raised by Bárány and Kalai, and independently Lew, we generalize Eckhoff's result to Cartesian products of convex sets in all dimensions. In particular, we prove that given $α\in (1-\frac{1}{t^d},1]$ and a finite family $\mathcal{F}$ of Cartesian products of convex sets $\prod_{i\in[t]}A_i$ in $\mathbb{R}^{td}$ with $A_i\subset \mathbb{R}^d$ if at least $α$-fraction of the $(d+1)$-tuples in $\mathcal{F}$ are intersecting then at least $(1-(t^d(1-α))^{1/(d+1)})$-fraction of sets in $\mathcal{F}$ are intersecting. This is a special case of a more general result on intersections of $d$-Leray complexes. We also provide a construction showing that our result on $d$-Leray complexes is optimal. Interestingly the extremal example is representable as a family of cartesian products of convex sets, implying the bound $α>1-\frac{1}{t^d}$ and the fraction $(1-(t^d(1-α))^{1/(d+1)})$ above are also best possible. The well-known optimal construction for fractional Helly theorem for convex sets in $\mathbb{R}^d$ does not have $(p,d+1)$-condition for sublinear $p$. Inspired by this we give constructions showing that, somewhat surprisingly, imposing additional $(p,d+1)$-condition has negligible effect on improving the quantitative bounds in neither the fractional Helly theorem for convex sets nor Cartesian products of convex sets. Our constructions offer a rich family of distinct extremal configurations for fractional Helly theorem, implying in a sense that the optimal bound is stable.

math.CO