SearcharxivSearch

arXiv subjects

Santanu Pal

Publications and source records attributed to Santanu Pal.

At least 19 recordsLinked to original sources

Which Tokens Need Context? A Reference-Based Analysis of Translation Responsibility Using Fertility and Entropy

When humans translate, not every word depends equally on the surrounding context. Some tokens, particularly function words like pronouns and auxiliaries, rely heavily on preceding or following sentences, while others, such as proper nouns, do not. Understanding this inherent context sensitivity is essential for evaluating whether machine translation systems use context in human-like ways. However, existing approaches to analysing context usage rely on discourse-specific test sets or model internals, making them narrow or model-dependent. We propose a post-hoc, model-agnostic framework to quantify context sensitivity at lexical and syntactic levels using two measures derived from word alignments: fertility (number of target tokens generated per source token) and entropy (stability of fertility patterns across contexts). Using reference translations for three language pairs (German $\leftrightarrow$ English, English $\rightarrow$ Hindi) under four context conditions, we show that context selectively redistributes generative responsibility from source to context tokens without altering overall fertility. Function words show the largest fertility reductions, while content words remain stable, suggesting that context resolves ambiguity rather than adding new information. Our framework provides a ground-truth characterisation of selective context usage in human translation, establishing a diagnostic baseline for evaluating machine translation models.

cs.CL

Beyond the Sentence: A Survey on Context-Aware Machine Translation with Large Language Models

Despite the popularity of the large language models (LLMs), their application to machine translation is relatively underexplored, especially in context-aware settings. This work presents a literature review of context-aware translation with LLMs. The existing works utilise prompting and fine-tuning approaches, with few focusing on automatic post-editing and creating translation agents for context-aware machine translation. We observed that the commercial LLMs (such as ChatGPT and Tower LLM) achieved better results than the open-source LLMs (such as Llama and Bloom LLMs), and prompt-based approaches serve as good baselines to assess the quality of translations. Finally, we present some interesting future directions to explore.

cs.CL

A Case Study on Context-Aware Neural Machine Translation with Multi-Task Learning

In document-level neural machine translation (DocNMT), multi-encoder approaches are common in encoding context and source sentences. Recent studies \cite{li-etal-2020-multi-encoder} have shown that the context encoder generates noise and makes the model robust to the choice of context. This paper further investigates this observation by explicitly modelling context encoding through multi-task learning (MTL) to make the model sensitive to the choice of context. We conduct experiments on cascade MTL architecture, which consists of one encoder and two decoders. Generation of the source from the context is considered an auxiliary task, and generation of the target from the source is the main task. We experimented with German--English language pairs on News, TED, and Europarl corpora. Evaluation results show that the proposed MTL approach performs better than concatenation-based and multi-encoder DocNMT models in low-resource settings and is sensitive to the choice of context. However, we observe that the MTL models are failing to generate the source from the context. These observations align with the previous studies, and this might suggest that the available document-level parallel corpora are not context-aware, and a robust sentence-level model can outperform the context-aware models.

cs.CL

TRAVID: An End-to-End Video Translation Framework

In today's globalized world, effective communication with people from diverse linguistic backgrounds has become increasingly crucial. While traditional methods of language translation, such as written text or voice-only translations, can accomplish the task, they often fail to capture the complete context and nuanced information conveyed through nonverbal cues like facial expressions and lip movements. In this paper, we present an end-to-end video translation system that not only translates spoken language but also synchronizes the translated speech with the lip movements of the speaker. Our system focuses on translating educational lectures in various Indian languages, and it is designed to be effective even in low-resource system settings. By incorporating lip movements that align with the target language and matching them with the speaker's voice using voice cloning techniques, our application offers an enhanced experience for students and users. This additional feature creates a more immersive and realistic learning environment, ultimately making the learning process more effective and engaging.

cs.CL

A Case Study on Context Encoding in Multi-Encoder based Document-Level Neural Machine Translation

Recent studies have shown that the multi-encoder models are agnostic to the choice of context, and the context encoder generates noise which helps improve the models in terms of BLEU score. In this paper, we further explore this idea by evaluating with context-aware pronoun translation test set by training multi-encoder models trained on three different context settings viz, previous two sentences, random two sentences, and a mix of both as context. Specifically, we evaluate the models on the ContraPro test set to study how different contexts affect pronoun translation accuracy. The results show that the model can perform well on the ContraPro test set even when the context is random. We also analyze the source representations to study whether the context encoder generates noise. Our analysis shows that the context encoder provides sufficient information to learn discourse-level information. Additionally, we observe that mixing the selected context (the previous two sentences in this case) and the random context is generally better than the other settings.

cs.CL

Quantum phases of ferromagnetically coupled dimers on Shastry-Sutherland lattice

The ground state (gs) of antiferromagnetically coupled dimers on the Shastry-Sutherland lattice (SSL) stabilizes many exotic phases and has been extensively studied. The gs properties of ferromagnetically coupled dimers on SSL are equally important but unexplored. In this model the exchange coupling along the $x$-axis ($J_x$) and $y$-axis ($J_y$) are ferromagnetic and the diagonal exchange coupling ($J$) is antiferromagnetic. In this work we explore the quantum phase diagram of ferromagnetically coupled dimer model numerically using density matrix renormalization group (DMRG) method. We note that in $J_x$-$J_y$ parameter space this model exhibits six interesting phases:(I) stripe $(0,π)$, (II) stripe $(π,0)$, (III) perfect dimer, (IV) $X$-spiral, (V) $Y$-spiral and (VI) ferromagnetic phase. Phase boundaries of these quantum phases are determined using the correlation functions and gs energies. We also notice the correlation length in this system is less than four lattice units in most of the parameter regimes. The non-collinear behaviour in $X$-spiral and $Y$-spiral phase and the dependence of pitch angles on model parameters are also studied.

cond-mat.str-el

Is Attention always needed? A Case Study on Language Identification from Speech

Language Identification (LID) is a crucial preliminary process in the field of Automatic Speech Recognition (ASR) that involves the identification of a spoken language from audio samples. Contemporary systems that can process speech in multiple languages require users to expressly designate one or more languages prior to utilization. The LID task assumes a significant role in scenarios where ASR systems are unable to comprehend the spoken language in multilingual settings, leading to unsuccessful speech recognition outcomes. The present study introduces convolutional recurrent neural network (CRNN) based LID, designed to operate on the Mel-frequency Cepstral Coefficient (MFCC) characteristics of audio samples. Furthermore, we replicate certain state-of-the-art methodologies, specifically the Convolutional Neural Network (CNN) and Attention-based Convolutional Recurrent Neural Network (CRNN with attention), and conduct a comparative analysis with our CRNN-based approach. We conducted comprehensive evaluations on thirteen distinct Indian languages and our model resulted in over 98\% classification accuracy. The LID model exhibits high-performance levels ranging from 97% to 100% for languages that are linguistically similar. The proposed LID model exhibits a high degree of extensibility to additional languages and demonstrates a strong resistance to noise, achieving 91.2% accuracy in a noisy setting when applied to a European Language (EU) dataset.

cs.LG

Colorful points in the XY regime of XXZ quantum magnets

In the XY regime of the XXZ Heisenberg model phase diagram, we demonstrate that the origin of magnetically ordered phases is influenced by the presence of solvable points with exact quantum coloring ground states featuring a quantum-classical correspondence. Using exact diagonalization and density matrix renormalization group calculations, for both the square and the triangular lattice magnets, we show that the ordered physics of the solvable points in the extreme XY regime, at $\frac{J_z}{J_\perp}=-1$ and $\frac{J_z}{J_\perp}=-\frac{1}{2}$ respectively with $J_\perp > 0$, adiabatically extends to the more isotropic regime $\frac{J_z}{J_\perp} \sim 1$. We highlight the projective structure of the coloring ground states to compute the correlators in fixed magnetization sectors which enables an understanding of the features in the static spin structure factors and correlation ratios. These findings are contrasted with an anisotropic generalization of the celebrated one-dimensional Majumdar-Ghosh model, which is also found to be (ground state) solvable. For this model, both exact dimer and three-coloring ground states exist at $\frac{J_z}{J_\perp}=-\frac{1}{2}$ but only the two dimer ground states survive for any $\frac{J_z}{J_\perp} > -\frac{1}{2}$.

cond-mat.str-el

Non-perturbative approach to quantum liquid ground states on geometrically frustrated Heisenberg antiferromagnets

We have formulated a twist operator argument for the geometrically frustrated quantum spin systems on the kagome and triangular lattices, thereby extending the application of the Lieb-Schultz-Mattis (LSM) and Oshikawa-Yamanaka-Affleck (OYA) theorems to these systems. The equivalent large gauge transformation for the geometrically frustrated lattice differs from that for non-frustrated systems due to the existence of multiple sublattices in the unit cell and non-orthogonal basis vectors. Our study for the $S=1/2$ kagome Heisenberg antiferromagnet at zero external magnetic field gives a criterion for the existence of a two-fold degenerate ground state with a finite excitation gap and fractionalized excitations. At finite field, we predict various plateaux at fractional magnetisation, in analogy with integer and fractional quantum Hall states of the primary sequence. These plateaux correspond to gapped quantum liquid ground states with a fixed number of singlets and spinons in the unit cell. A similar analysis for the triangular lattice predicts a single fractional magnetization plateau at $1/3$. Our results are in broad agreement with numerical and experimental studies.

cond-mat.str-el

The Transference Architecture for Automatic Post-Editing

In automatic post-editing (APE) it makes sense to condition post-editing (pe) decisions on both the source (src) and the machine translated text (mt) as input. This has led to multi-source encoder based APE approaches. A research challenge now is the search for architectures that best support the capture, preparation and provision of src and mt information and its integration with pe decisions. In this paper we present a new multi-source APE model, called transference. Unlike previous approaches, it (i) uses a transformer encoder block for src, (ii) followed by a decoder block, but without masking for self-attention on mt, which effectively acts as second encoder combining src -> mt, and (iii) feeds this representation into a final decoder block generating pe. Our model outperforms the state-of-the-art by 1 BLEU point on the WMT 2016, 2017, and 2018 English--German APE shared tasks (PBSMT and NMT). We further investigate the importance of our newly introduced second encoder and find that a too small amount of layers does hurt the performance, while reducing the number of layers of the decoder does not matter much.

cs.CL

UDS--DFKI Submission to the WMT2019 Similar Language Translation Shared Task

In this paper we present the UDS-DFKI system submitted to the Similar Language Translation shared task at WMT 2019. The first edition of this shared task featured data from three pairs of similar languages: Czech and Polish, Hindi and Nepali, and Portuguese and Spanish. Participants could choose to participate in any of these three tracks and submit system outputs in any translation direction. We report the results obtained by our system in translating from Czech to Polish and comment on the impact of out-of-domain test data in the performance of our system. UDS-DFKI achieved competitive performance ranking second among ten teams in Czech to Polish translation.

cs.CL

Improving CAT Tools in the Translation Workflow: New Approaches and Evaluation

This paper describes strategies to improve an existing web-based computer-aided translation (CAT) tool entitled CATaLog Online. CATaLog Online provides a post-editing environment with simple yet helpful project management tools. It offers translation suggestions from translation memories (TM), machine translation (MT), and automatic post-editing (APE) and records detailed logs of post-editing activities. To test the new approaches proposed in this paper, we carried out a user study on an English--German translation task using CATaLog Online. User feedback revealed that the users preferred using CATaLog Online over existing CAT tools in some respects, especially by selecting the output of the MT system and taking advantage of the color scheme for TM suggestions.

cs.CL

Search for stabilizing effects of $\bm{Z=82}$ shell closure against fission

Presence of closed proton and/or neutron shells causes deviation from macroscopic properties of nuclei which are understood in terms of the liquid drop model. It is important to investigate experimentally the stabilizing effects of shell closure, if any, against fission. This work aims to investigate probable effects of proton shell ($Z = 82$) closure in the compound nucleus, in enhancing survival probability of the evaporation residues formed in heavy ion-induced fusion-fission reactions. Evaporation residue cross sections have been measured for the reactions $^{19}$F+$^{180}$Hf, $^{19}$F+$^{181}$Ta and $^{19}$F+$^{182}$W from $\simeq9\%$ below to $\simeq42\%$ above the Coulomb barrier, leading to formation of compound nuclei with same number of neutrons ($N = 118$) but different number of protons across $Z = 82$. Measured excitation functions have been compared with statistical model calculation, in which reduced dissipation coefficient is the only adjustable parameter. Evaporation residue cross section, normalized by capture cross section, is found to decrease gradually with increasing fissility of the compound nucleus. Measured evaporation residue cross sections require inclusion of nuclear viscosity in the model calculations. Reduced dissipation coefficient in the range of 1\textendash3 $\times$ $10^{21}$ s$^{-1}$ reproduces the data quite well. No abrupt enhancement of evaporation residue cross sections has been observed in the reaction forming compound nucleus with $Z = 82$. Thus, this work does not find enhanced stabilizing effects of $Z = 82$ shell closure against fission in the compound nucleus. One may attempt to measure cross sections of individual exit channels for further confirmation of our observation.

nucl-ex

Integrating Artificial and Human Intelligence for Efficient Translation

Current advances in machine translation increase the need for translators to switch from traditional translation to post-editing of machine-translated text, a process that saves time and improves quality. Human and artificial intelligence need to be integrated in an efficient way to leverage the advantages of both for the translation task. This paper outlines approaches at this boundary of AI and HCI and discusses open research questions to further advance the field.

cs.HC

Magnetisation plateaux of the quantum pyrochlore Heisenberg antiferromagnet

We predict magnetisation plateaux ground states for $S=1/2$ Heisenberg antiferromagnets on pyrochlore lattices by formulating arguments based on gauge and spin-parity transformations. We derive a twist operator appropriate to the pyrochlore lattice, and show that it is equivalent to a large gauge transformation. Invariance under this large gauge transformation indicates the sensitivity of the ground state to changes in boundary conditions. This leads to the formulation of an Oshikawa-Yamanaka-Affleck (OYA)-like criterion at finite external magnetic field, enabling the prediction of plateaux in the magnetisation versus field diagram. We also develop an analysis based on the spin-parity operator, leading to a condition from which identical predictions are obtained of magnetisation plateaux ground states. Both analyses are based on the non-local nature of the transformations, and rely only on the symmetries of the Hamiltonian. This suggests that the plateaux ground states can possess properties arising from non-local entanglement between the spins. We also demonstrate that while a spin-lattice coupling stabilises plateaux in a system of quantum spins with antiferromagnetic exchange, it can compete with weak ferromagnetic spin exchange in leading to frustration-induced magnetisation plateaux.

cond-mat.str-el

Correlated spin liquids in the quantum kagome antiferromagnet at finite field: a renormalisation group analysis

We analyse the antiferromagnetic spin-$1/2$ XXZ model on the kagome lattice at finite external magnetic field with the help of a nonperturbative zero-temperature renormalization group (RG) technique. Following the work of Kumar \emph{et al} (Phys. Rev. B {\bf 90}, 174409 (2014)), we use a Jordan-Wigner transformation to map the spin problem into one of spinless fermions (spinons) in the presence of a statistical gauge field, and with nearest-neighbour interactions. While the work of Kumar \emph{et al} was confined mostly to the plateau at $1/3$-filling (magnetisation per site) in the XY regime, we analyse the role of inter-spinon interactions in shaping the phases around this plateau in the entire XXZ model. The RG phase diagram obtained contains three spin liquid phases whose position is determined as a function of the exchange anisotropy and the energy scale for fluctuations arising from spinon scattering. Two of these spins liquids are topologically ordered states of matter with gapped, degenerate states on the torus. The gap for one of these phases corresponds to the one-spinon band gap of the Azbel-Hofstadter spectrum for the XY part of the Hamiltonian, while the other arises from two-spinon interactions. The Heisenberg point of this problem is found to lie within the interaction gapped spin liquid phase, in broad agreement with a recent experimental finding. The third phase is an algebraic spin liquid with a gapless Dirac spectrum for spinon excitations, and possess properties that show departures from the Fermi liquid paradigm. The three phase boundaries correspond to critical theories, and meet at a $SU(2)$-symmetric multicritical point. This special critical point agrees well with the gap-closing transition point predicted by Kumar \emph{et al}. We discuss the relevance of our findings to various recent experiments, as well as results obtained from other theoretical analyses.

cond-mat.str-el

Alleviating the inconsistencies in modelling decay of fissile compound nuclei

This work attempts to overcome the existing inconsistencies in modelling decay of fissile nucleus by inclusion of important physical effects in the model and through a systematic analysis of a large set of data over a wide range of CN mass (ACN). The model includes shell effect in the level density (LD) parameter, shell correction in the fission barrier, effect of the orientation degree of freedom of the CN spin (Kor), collective enhancement of level density (CELD) and dissipation in fission. Input parameters are not tuned to reproduce observables from specific reaction(s) and the reduced dissipation coefficient is treated as the only adjustable parameter. Calculated evaporation residue (ER) cross sections, fission cross sections and particle, i.e. neutron, proton and alpha-particle, multiplicities are compared with data covering ACN = 156-248. The model produces reasonable fits to ER and fission excitation functions for all the reactions considered in this work. Pre-scission neutron multiplicities are underestimated by the calculation beyond ACN~200. An increasingly higher value of pre-saddle dissipation strength is required to reproduce the data with increasing ACN. Proton and alpha-particle multiplicities, measured in coincidence with both ERs and fission fragments, are in qualitative agreement with model predictions. The present work mitigates the existing inconsistencies in modelling statistical decay of the fissile CN to a large extent.

nucl-th

Discriminating between Indo-Aryan Languages Using SVM Ensembles

In this paper we present a system based on SVM ensembles trained on characters and words to discriminate between five similar languages of the Indo-Aryan family: Hindi, Braj Bhasha, Awadhi, Bhojpuri, and Magahi. We investigate the performance of individual features and combine the output of single classifiers to maximize performance. The system competed in the Indo-Aryan Language Identification (ILI) shared task organized within the VarDial Evaluation Campaign 2018. Our best entry in the competition, named ILIdentification, scored 88:95% F1 score and it was ranked 3rd out of 8 teams.

cs.CL