SearcharxivSearch

arXiv subjects

Mintaek Oh

Publications and source records attributed to Mintaek Oh.

9 recordsLinked to original sources

CanonNav: Disentangling Navigation Behavior from Camera Geometry in Cross-Platform Visual Navigation

While visual navigation has advanced through imitation learning from cross-platform demonstrations, fully leveraging such data remains challenging. First, directly learning from image-trajectory pairs entangles navigation behavior with platform-dependent camera geometry. This hinders consistent learning by forcing the policy to implicitly infer camera geometry from visual observations, an inherently ill-posed problem. Second, imitation learning from demonstrated trajectories captures the expert's chosen motion but leaves the intermediate decisions underlying that motion implicit. To address these issues, we propose CanonNav, a visual navigation framework that disentangles navigation behavior from camera geometry and incorporates complementary planning supervision into learning from cross-platform demonstrations. CanonNav introduces camera geometry canonicalization, which transforms visual observations and trajectories into a camera-consistent representation space. Building on this representation, we derive safety and local-progress supervision using pseudo-labels from an offline traversability estimator. Safety supervision penalizes unsafe trajectories, while local-progress supervision guides where the robot should advance. Experiments across diverse camera configurations and environments show that, despite using only RGB at inference, CanonNav consistently outperforms RGB-based baselines and even surpasses RGB-D-based methods in challenging scenarios.

cs.RO

Scalable and Convergent Generalized Power Iteration Precoding for Massive MIMO Systems

In massive multiple-input multiple-output (MIMO) systems, achieving high spectral efficiency (SE) often requires advanced precoding algorithms whose complexity scales rapidly with the number of antennas, limiting practical deployment. In this paper, we develop a scalable and computationally efficient generalized power iteration precoding (GPIP) framework for massive MIMO systems under both perfect and imperfect channel state information at the transmitter (CSIT). By exploiting the low-dimensional subspace property of optimal precoders, we reformulate the high-dimensional beamforming problem into a lower-dimensional weight optimization that scales with the number of users rather than antennas. We further extend this framework to the imperfect CSIT scenario by showing that stationary solutions reside in a combined subspace spanned by the estimated channel and error covariance matrices, enabling a robust design via low-rank approximation. To reduce computational cost, we leverage the Sherman-Morrison formula to simplify matrix inversions. Moreover, interpreting the GPIP update as a projected preconditioned gradient ascent method, we establish convergence guarantees and develop a stable and monotonic algorithm using a backtracking line search. Numerical results demonstrate that the proposed methods achieve the highest SE performance compared to state-of-the-art linear precoders with significantly reduced complexity and convergence, highlighting their suitability for large-scale MIMO systems.

eess.SP

Hybrid Precoding Revisited: Low-Dimensional Subspace Perspective for MU-MIMO Systems

This letter presents a low-complexity hybrid precoding framework for multiuser multiple-input multiple-output (MIMO) systems by leveraging a low-dimensional subspace property. Under the low-dimensional subspace perspective, we first identify an unconstrained optimal radio-frequency (RF) precoder. We then optimize a hybrid precoder via a reduced-complexity precoding method. We further extend the proposed framework to (i) a dynamic-subarray antenna partitioning algorithm that adaptively allocates subsets of antennas associated with RF chains, and (ii) a channel covariance-based approach to exploit statistical channel state information at a transmitter (CSIT), ensuring robustness with partial CSIT. Simulations validate that our proposed algorithms achieve superior performance while significantly reducing complexity compared to existing methods.

eess.SP

Language as Cost: Proactive Hazard Mapping using VLM for Robot Navigation

Robots operating in human-centric or hazardous environments must proactively anticipate and mitigate dangers beyond basic obstacle detection. Traditional navigation systems often depend on static maps, which struggle to account for dynamic risks, such as a person emerging from a suddenly opening door. As a result, these systems tend to be reactive rather than anticipatory when handling dynamic hazards. Recent advancements in pre-trained large language models and vision-language models (VLMs) create new opportunities for proactive hazard avoidance. In this work, we propose a zero-shot language-as-cost mapping framework that leverages VLMs to interpret visual scenes, assess potential dynamic risks, and assign risk-aware navigation costs preemptively, enabling robots to anticipate hazards before they materialize. By integrating this language-based cost map with a geometric obstacle map, the robot not only identifies existing obstacles but also anticipates and proactively plans around potential hazards arising from environmental dynamics. Experiments in simulated and diverse dynamic environments demonstrate that the proposed method significantly improves navigation success rates and reduces hazard encounters, compared to reactive baseline planners. Code and supplementary materials are available at https://github.com/Taekmino/LaC.

cs.RO

Scalable Beamforming Design for Multi-RIS-Aided MU-MIMO Systems with Imperfect CSIT

This paper presents a scalable beamforming design for maximizing the spectral efficiency (SE) of multi-reconfigurable intelligent surface (RIS)-aided communications through joint optimization of the precoder and RIS phase shifts in multi-user multiple-input multiple-output (MU-MIMO) systems under imperfect channel state information at the transmitter (CSIT). To address key challenges of the joint optimization problem, we first decompose it into two subproblems by deriving a proper lower bound. We then leverage a generalized power iteration (GPI) approach to identify a superior local optimal precoding solution. We further extend this approach to the RIS design using regularization; we set a RIS regularization function to efficiently handle the unit-modulus constraints, and also find the superior local optimal solution for RIS phase shifts under the GPI-based optimization framework. Subsequently, we propose an alternating optimization method. Our proposed algorithm offers scalable multi-RIS beamforming in terms of computational complexity that scales linearly with the number of RISs, while achieving superior performance. We further reduce the complexity with respect to the number of RIS elements by using diagonal approximation of the channel error covariance and avoiding direct matrix inversion. Simulations validate the proposed algorithm in terms of both the sum SE performance and the scalability.

eess.SP

E2Map: Experience-and-Emotion Map for Self-Reflective Robot Navigation with Language Models

Large language models (LLMs) have shown significant potential in guiding embodied agents to execute language instructions across a range of tasks, including robotic manipulation and navigation. However, existing methods are primarily designed for static environments and do not leverage the agent's own experiences to refine its initial plans. Given that real-world environments are inherently stochastic, initial plans based solely on LLMs' general knowledge may fail to achieve their objectives, unlike in static scenarios. To address this limitation, this study introduces the Experience-and-Emotion Map (E2Map), which integrates not only LLM knowledge but also the agent's real-world experiences, drawing inspiration from human emotional responses. The proposed methodology enables one-shot behavior adjustments by updating the E2Map based on the agent's experiences. Our evaluation in stochastic navigation environments, including both simulations and real-world scenarios, demonstrates that the proposed method significantly enhances performance in stochastic environments compared to existing LLM-based approaches. Code and supplementary materials are available at https://e2map.github.io/.

cs.RO

Full-Duplex Multiuser MISO Under Coarse Quantization: Per-Antenna SQNR Analysis and Beamforming Design

We investigate full-duplex (FD) multi-user multiple input single-output systems with coarse quantization, aiming to characterize the impact of employing low-resolution analog-to-digital converters (ADCs) on self-interference (SI) and to develop a quantization- and SI-aware beamforming method that alleviates quantization-induced performance degradation in the FD systems. We first present an analysis on the perantenna signal-to-quantization noise ratio for conventional linear beamformers to provide the desired range of the number of analog-to-digital converter (ADC) bits, providing system insights for reliable FD operation in regard to the ADC resolution and beamforming strategy. Motivated by the insights, we then propose an SI-aware beamforming method that mitigates residual SI and quantization distortion. The resulting spectral efficiency (SE) maximization problem is decomposed into two tractable subproblems solved via alternating optimization: precoder and combiner design. The precoder optimization is formulated as a generalized eigenvalue problem, where the dominant eigenvector yields the best stationary solution through power iteration, while the combiner is derived as a quantization-aware minimum meansquared error (MMSE) filter. Numerical studies show that the number of required ADC bits with the proposed beamforming falls within the derived theoretical range while achieving the highest SE compared to benchmarks.

cs.IT

Joint Optimization for Secure and Reliable Communications in Finite Blocklength Regime

To realize ultra-reliable low latency communications with high spectral efficiency and security, we investigate a joint optimization problem for downlink communications with multiple users and eavesdroppers in the finite blocklength (FBL) regime. We formulate a multi-objective optimization problem to maximize a sum secrecy rate by developing a secure precoder and to minimize a maximum error probability and information leakage rate. The main challenges arise from the complicated multi-objective problem, non-tractable back-off factors from the FBL assumption, non-convexity and non-smoothness of the secrecy rate, and the intertwined optimization variables. To address these challenges, we adopt an alternating optimization approach by decomposing the problem into two phases: secure precoding design, and maximum error probability and information leakage rate minimization. In the first phase, we obtain a lower bound of the secrecy rate and derive a first-order Karush-Kuhn-Tucker (KKT) condition to identify local optimal solutions with respect to the precoders. Interpreting the condition as a generalized eigenvalue problem, we solve the problem by using a power iteration-based method. In the second phase, we adopt a weighted-sum approach and derive KKT conditions in terms of the error probabilities and leakage rates for given precoders. Simulations validate the proposed algorithm.

cs.IT

Joint Precoding and Artificial Noise Design for MU-MIMO Wiretap Channels

Secure precoding superimposed with artificial noise (AN) is a promising transmission technique to improve security by harnessing the superposition nature of the wireless medium. However, finding a jointly optimal precoding and AN structure is very challenging in downlink multi-user multiple-input multiple-output (MU-MIMO) wiretap channels with multiple eavesdroppers. The major challenge in maximizing the secrecy rate arises from the non-convexity and non-smoothness of the rate function. Traditionally, an alternating optimization framework that identifies beamforming vectors and AN covariance matrix has been adopted; yet this alternating approach has limitations in maximizing the secrecy rate. In this paper, we put forth a novel secure precoding algorithm that jointly and simultaneously optimizes the beams and AN covariance matrix for maximizing the secrecy rate when a transmitter has either perfect or partial channel knowledge of eavesdroppers. To this end, we first establish an approximate secrecy rate in a smooth function. Then, we derive the first-order optimality condition in the form of the nonlinear eigenvalue problem (NEP). We present a computationally efficient algorithm to identify the principal eigenvector of the NEP as a suboptimal solution for secure precoding. Simulations demonstrate that the proposed methods improve secrecy rate significantly compared to the existing secure precoding methods.

cs.IT