SearcharxivSearch

arXiv subjects

Shinya Wada

Publications and source records attributed to Shinya Wada.

11 recordsLinked to original sources

xMTrans: Temporal Attentive Cross-Modality Fusion Transformer for Long-Term Traffic Prediction

Traffic predictions play a crucial role in intelligent transportation systems. The rapid development of IoT devices allows us to collect different kinds of data with high correlations to traffic predictions, fostering the development of efficient multi-modal traffic prediction models. Until now, there are few studies focusing on utilizing advantages of multi-modal data for traffic predictions. In this paper, we introduce a novel temporal attentive cross-modality transformer model for long-term traffic predictions, namely xMTrans, with capability of exploring the temporal correlations between the data of two modalities: one target modality (for prediction, e.g., traffic congestion) and one support modality (e.g., people flow). We conducted extensive experiments to evaluate our proposed model on traffic congestion and taxi demand predictions using real-world datasets. The results showed the superiority of xMTrans against recent state-of-the-art methods on long-term traffic predictions. In addition, we also conducted a comprehensive ablation study to further analyze the effectiveness of each module in xMTrans.

cs.LG

VideoAdviser: Video Knowledge Distillation for Multimodal Transfer Learning

Multimodal transfer learning aims to transform pretrained representations of diverse modalities into a common domain space for effective multimodal fusion. However, conventional systems are typically built on the assumption that all modalities exist, and the lack of modalities always leads to poor inference performance. Furthermore, extracting pretrained embeddings for all modalities is computationally inefficient for inference. In this work, to achieve high efficiency-performance multimodal transfer learning, we propose VideoAdviser, a video knowledge distillation method to transfer multimodal knowledge of video-enhanced prompts from a multimodal fundamental model (teacher) to a specific modal fundamental model (student). With an intuition that the best learning performance comes with professional advisers and smart students, we use a CLIP-based teacher model to provide expressive multimodal knowledge supervision signals to a RoBERTa-based student model via optimizing a step-distillation objective loss -- first step: the teacher distills multimodal knowledge of video-enhanced prompts from classification logits to a regression logit -- second step: the multimodal knowledge is distilled from the regression logit of the teacher to the student. We evaluate our method in two challenging multimodal tasks: video-level sentiment analysis (MOSI and MOSEI datasets) and audio-visual retrieval (VEGAS dataset). The student (requiring only the text modality as input) achieves an MAE score improvement of up to 12.3% for MOSI and MOSEI. Our method further enhances the state-of-the-art method by 3.4% mAP score for VEGAS without additional computations for inference. These results suggest the strengths of our method for achieving high efficiency-performance multimodal transfer learning.

cs.CV

VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question Answering

Visual question answering (VQA) requires systems to perform concept-level reasoning by unifying unstructured (e.g., the context in question and answer; "QA context") and structured (e.g., knowledge graph for the QA context and scene; "concept graph") multimodal knowledge. Existing works typically combine a scene graph and a concept graph of the scene by connecting corresponding visual nodes and concept nodes, then incorporate the QA context representation to perform question answering. However, these methods only perform a unidirectional fusion from unstructured knowledge to structured knowledge, limiting their potential to capture joint reasoning over the heterogeneous modalities of knowledge. To perform more expressive reasoning, we propose VQA-GNN, a new VQA model that performs bidirectional fusion between unstructured and structured multimodal knowledge to obtain unified knowledge representations. Specifically, we inter-connect the scene graph and the concept graph through a super node that represents the QA context, and introduce a new multimodal GNN technique to perform inter-modal message passing for reasoning that mitigates representational gaps between modalities. On two challenging VQA tasks (VCR and GQA), our method outperforms strong baseline VQA methods by 3.2% on VCR (Q-AR) and 4.6% on GQA, suggesting its strength in performing concept-level reasoning. Ablation studies further demonstrate the efficacy of the bidirectional fusion and multimodal GNN method in unifying unstructured and structured multimodal knowledge.

cs.CV

Valley Views: Instantons, Large Order Behaviors, and Supersymmetry

The elucidation of the properties of the instantons in the topologically trivial sector has been a long-standing puzzle. Here we claim that the properties can be summarized in terms of the geometrical structure in the configuration space, the valley. The evidence for this claim is presented in various ways. The conventional perturbation theory and the non-perturbative calculation are unified, and the ambiguity of the Borel transform of the perturbation series is removed. A `proof' of Bogomolny's ``trick'' is presented, which enables us to go beyond the dilute-gas approximation. The prediction of the large order behavior of the perturbation theory is confirmed by explicit calculations, in some cases to the 478-th order. A new type of supersymmetry is found as a by-product, and our result is shown to be consistent with the non-renormalization theorem. The prediction of the energy levels is confirmed with numerical solutions of the Schrödinger equation.

hep-th

Valleys in Quantum Mechanics

Conventionally, perturbative and non-perturbative calculations are performed independently. In this paper, valleys in the configuration space in quantum mechanics are investigated as a way to treat them in a unified manner. All the known results of the interplay of them are reproduced naturally. The prescription for separating the non-perturbative contribution from the perturbative is given in terms of the analytic continuation of the valley parameter. Our method is illustrated on a new series of examples with the asymmetric double-well potential. We obtain the non-perturbative part explicitly, which leads to the prediction of the large order behavior of the perturbative series. We calculate the first 200 perturbative coefficients for a wide range of parameters and confirm the agreement with the prediction of the valley method.

quant-ph

Recent Developments in the Theory of Tunneling

Path-integral approach in imaginary and complex time has been proven successful in treating the tunneling phenomena in quantum mechanics and quantum field theories. Latest developments in this field, the proper valley method in imaginary time, its application to various quantum systems, complex time formalism, asympton theory for the large order analysis of the perturbation theory, are reviewed in a self-contained manner.

hep-th

Multi-instanton calculus in N=2 supersymmetric QCD

Microscopic tests of the exact results are performed in N=2 SU(2) supersymmetric QCD. We construct the multi-instanton solution in N=2 supersymmetric QCD and calculate the two-instanton contribution ${\cal F}_2$ to the prepotential ${\cal F}$ explicitly. For $N_f=1,2$, instanton calculus agrees with the prediction of the exact results, however, for $N_f=3$, we find a discrepancy between them.

hep-th

Fake Instability in the Euclidean Formalism

We study the path-integral formalism in the imaginary-time to show its validity in a case with a metastable ground state. The well-known method based on the bounce solution leads to the imaginary part of the energy even for a state that is only metastable and has a simple oscillating behavior instead of decaying. Although this has been argued to be the failure of the Euclidean formalism, we show that proper account of the global structure of the path-space leads to a valid expression for the energy spectrum, without the imaginary part. For this purpose we use the proper valley method to find a new type of instanton-like configuration, the ``valley instantons''. Although valley instantons are not the solutions of equation of motion, they have dominant contribution to the functional integration. A dilute-gas approximation for the valley instantons is shown to lead to the energy formula. This method extends the well-known imaginary-time formalism so that it can take into account the global behavior of the theory.

hep-th

Valley Instanton versus Constrained Instanton

Based on the new valley equation, we propose the most plausible method for constructing instanton-like configurations in the theory where the presence of a mass scale prevents the existence of the classical solution with a finite radius. We call the resulting instanton-like configuration as valley instanton. The detail comparison between the valley instanton and the constrained instanton in $ϕ^4$ theory and the gauge-Higgs system are carried out. For instanton-like configurations with large radii, there appear remarkable differences between them. These differences are essential in calculating the baryon number violating processes with multi bosons.

hep-th

Valley Instanton in the Gauge-Higgs System

The instanton configuration in the SU(2)-gauge system with a Higgs doublet is constructed by using the new valley method. This method defines the configuration by an extension of the field equation and allows the exact conversion of the quasi-zero eigenmode to a collective coordinate. It does not require ad-hoc constraints used in the current constrained instanton method and provides a better mathematical formalism than the constrained instanton method. The resulting instanton, which we call ``valley instanton'', is shown to have desirable behaviors. The result of the numerical investigation is also presented.

hep-th

Bounce in Valley: Study of the extended structures from thick-wall to thin-wall vacuum bubbles

The valley structure associated with quantum meta-stability is examined. It is defined by the new valley equation, which enables consistent evaluation of the imaginary-time path-integral. We study the structure of this new valley equation and solve these equations numerically. The valley is shown to contain the bounce solution, as well as other bubble structures. We find that even when the bubble solution has thick wall, the outer region of the valley is made of large-radius, thin-wall bubble, which interior is occupied by the true-vacuum. Smaller size bubbles, which contribute to decay at higher energies, are also identified.

hep-th