SearcharxivSearch

arXiv subjects

Haocheng Ren

Publications and source records attributed to Haocheng Ren.

3 recordsLinked to original sources

Compact Representation of Mipmapped SVBRDFs via Shared Gaussians

Spatially-varying BRDFs (SVBRDFs) are central to material representation in computer graphics, but their high-resolution, multi-channel, mipmapped textures impose a substantial storage burden. Existing compression methods face a fundamental trade-off: block-based compression provides random access and hardware-friendly decoding but exploits redundancy only within local blocks; image codecs offer strong rate-distortion performance but are not designed for direct real-time texture access; and neural texture compression achieves high compression ratios but requires neural inference during decoding, which introduces additional runtime overhead, especially on mobile platforms. We present Gaussian Texture Compression (GTC), a compact 2D Gaussian-based representation for mipmapped SVBRDF texture stacks that delivers high-quality compression with flexible rate-distortion trade-offs. Our method is based on a key observation that there are two dominant sources of redundancy in such data: across mip levels and across material maps. Both share a common underlying structure: the same spatial support is reused, with only level- or map-specific information attached. This property naturally suits 2D Gaussians, since each Gaussian explicitly separates its spatial footprint from the values it carries, allowing the footprint to be shared while the values vary per level and per map. Building on this property, GTC shares Gaussians along both redundancy dimensions and is trained via a progressive optimization pipeline. Experiments show that GTC achieves higher reconstruction quality and lower memory usage than ASTC, the industry-standard GPU texture compression format, while supporting random-access, non-neural decoding suitable for real-time rendering.

cs.GR

Rubikon: Intelligent Tutoring for Rubik's Cube Learning Through AR-enabled Physical Task Reconfiguration

Learning to solve a Rubik's Cube requires the learners to repeatedly practice a skill component, e.g., identifying a misplaced square and putting it back. However, for 3D physical tasks such as this, generating sufficient repeated practice opportunities for learners can be challenging, in part because it is difficult for novices to reconfigure the physical object to specific states. We propose Rubikon, an intelligent tutoring system for learning to solve the Rubik's Cube. Rubikon reduces the necessity for repeated manual configurations of the Rubik's Cube without compromising the tactile experience of handling a physical cube. The foundational design of Rubikon is an AR setup, where learners manipulate a physical cube while seeing an AR-rendered cube on a display. Rubikon automatically generates configurations of the Rubik's Cube to target learners' weaknesses and help them exercise diverse knowledge components. In a between-subjects experiment, we showed that Rubikon learners scored 25% higher on a post-test compared to baselines.

cs.HC

MINERVAS: Massive INterior EnviRonments VirtuAl Synthesis

With the rapid development of data-driven techniques, data has played an essential role in various computer vision tasks. Many realistic and synthetic datasets have been proposed to address different problems. However, there are lots of unresolved challenges: (1) the creation of dataset is usually a tedious process with manual annotations, (2) most datasets are only designed for a single specific task, (3) the modification or randomization of the 3D scene is difficult, and (4) the release of commercial 3D data may encounter copyright issue. This paper presents MINERVAS, a Massive INterior EnviRonments VirtuAl Synthesis system, to facilitate the 3D scene modification and the 2D image synthesis for various vision tasks. In particular, we design a programmable pipeline with Domain-Specific Language, allowing users to (1) select scenes from the commercial indoor scene database, (2) synthesize scenes for different tasks with customized rules, and (3) render various imagery data, such as visual color, geometric structures, semantic label. Our system eases the difficulty of customizing massive numbers of scenes for different tasks and relieves users from manipulating fine-grained scene configurations by providing user-controllable randomness using multi-level samplers. Most importantly, it empowers users to access commercial scene databases with millions of indoor scenes and protects the copyright of core data assets, e.g., 3D CAD models. We demonstrate the validity and flexibility of our system by using our synthesized data to improve the performance on different kinds of computer vision tasks.

cs.CV