Searcharxiv⌕ Search

arXiv subjects

Chaitanya Lakhchaura

Publications and source records attributed to Chaitanya Lakhchaura.

2 recordsLinked to original sources

NutriVision: Ingredient-Conditioned Fusion and Prediction for Single-Image Food Nutrition Estimation

Nutrition estimation is a fundamental task in consumer diet tracking, clinical dietetics, chronic disease management, sports and hospital nutrition, and broader food computing systems. The existing approaches have progressed along two largely separate axes, vision models that rely on calibrated RGB-depth captures and ingredient-aware methods that use textual cues but use limited multimodal fusion. We introduce NutriVision, an end-to-end framework that leverages visual geometry and ingredient semantics to estimate calories, mass, fat content, carbohydrates, and protein from a single RGB image and an optional ingredient list. It obtains the unavailable depth modality using DepthAnything-V3 and encodes ingredient descriptions using CLIP. It integrates three complementary mechanisms: (1) an \emph{Ingredient-Conditioned Frequency-Aligned Fusion Module (IC-FAFM)}, which uses textual guidance to reweight and align RGB-depth frequency components; (2) an \emph{Ingredient-Aware Mask-based Prediction Head (IA-MPH)}, whose gating and channel masks are conditioned on food identity; and (3) modality-specific \emph{Internal Semantic Modeling (ISM)} blocks. On the Nutrition5k dataset, NutriVision achieves a mean PMAE of $\mathbf{13.60\pm0.10\%}$, outperforming our IGSMNet implementation by $0.89$ percentage points and OmniFood8k by $2.90$ percentage points (both $p<0.001$). The module-level ablations identify the ingredient-aware prediction head as the primary architectural contributor, improving mean PMAE by $1.50\pm0.17$ percentage points ($p<0.001$). These results demonstrate that ingredient-conditioned prediction and frequency-aware RGB-depth fusion provide measurable gains for single-image nutrient estimation. More broadly, NutriVision offers a practical route toward nutrition-assessment systems that exploit geometric and semantic cues without requiring specialized depth-sensing hardware

cs.CV↗

dKFD: Phase-Structured Evidence Allocation for Fixed-Budget Localized Event Understanding

Sparse video understanding often requires selecting a small set of visual evidence under a fixed frame budget. Most sparse selectors allocate this budget globally, allowing all frames to compete with one another. For temporally localized events, this can be a poor inductive bias: useful evidence is often distributed across pre-event context, the event itself, and post-event consequences. We study fixed-budget evidence allocation for localized event videos and show that globally competitive Top-$K$ selectors can preserve recognition and grounding while producing unstable event evidence. On DoTA Video Anomaly Recognition, Global Top-$K$ obtains competitive recognition and temporal grounding, but low selector-event alignment at $K=12$ (Frame AUC $51.5 \pm 10.8$). We propose dKFD, a phase-structured differentiable selector that reserves evidence capacity across pre-event, event, and post-event phases after full-sequence temporal encoding. Under matched-budget multi-seed evaluation, dKFD improves Frame AUC by $+30.97$ over a matched Global Top-$K$ selector at $K=12$ ($p<0.01$), while yielding modest but statistically significant recognition gains and comparable temporal grounding. Mechanism ablations show that phase supervision is load-bearing: removing it reduces Frame AUC to $41.1 \pm 12.0$ even when phase-partitioned budgets are retained. Downstream diagnostics on VRU-Accident show consistent gains over learned Global Top-$K$ across VLM families, while dense captioning reveals a boundary condition where uniform sampling remains competitive. These results support phase-structured allocation as a controlled fixed-budget approach for event-centric sparse evidence selection, not as a universal video summarization strategy.

cs.CV↗