SearcharxivSearch

arXiv subjects

Binnan Liu

Publications and source records attributed to Binnan Liu.

2 recordsLinked to original sources

TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and apply it to a new grid. Looped visual reasoners refine predictions over multiple iterations, but conventional training constrains only the final output, leaving intermediate refinements unconstrained. We propose that these refinements should instead follow the transformation step by step. We introduce TraceViT, a looped visual reasoner trained with semantically monotonic transformation chains. We obtain these chains by rewriting and verifying programmatic task implementations, decomposing each solution into intermediate grid states. Each iteration is grounded by a task reference derived from the few-shot demonstrations and an object workspace representing the current grid state. Because these chains may differ in length from the loop, soft trace alignment enforces only their ordering, letting the model allocate iterations freely. TraceViT achieves 67.8% pass@2 on ARC-AGI-1 and 24.3% on ARC-AGI-2. Controlled ablations on ARC-AGI-1 show that trace supervision becomes beneficial only when paired with grounding. Code and data will be available at https://github.com/LiuBinnan/TraceViT.

cs.CV

RefRetouch: Personalized Image Retouching without Test-time Fine-tuning

Personalized image retouching aims to adapt retouching styles of individual users from reference examples, but existing methods often require user-specific fine-tuning or fail to generalize effectively. To address these challenges, we introduce \textbf{RefRetouch}, a general framework for personalized image retouching that instantly adapts to user retouching styles without any test-time fine-tuning. It employs an \textit{asymmetric auto-encoder} to encode the retouching style from paired examples into a content disentangled latent representation that enables faithful transfer of the retouching style to new images. To adaptively apply the encoded retouching style to new images, we further propose \textit{retrieval-augmented retouching} (RAR), which retrieves and aggregates style latents from reference pairs most similar in content to the query image. With these components, \textbf{RefRetouch} enables superior and generic content-aware retouching personalization across diverse scenarios, including single-reference, multi-reference, and mixed-style settings, while also generalizing out of the box to photorealistic style transfer.

cs.GR