Edit Distances and Their Applications to Downstream Tasks in Research and Commercial Contexts

  • 2024-10-08 11:21:22
  • FĂ©lix do Carmo, Diptesh Kanojia
  • 0

Abstract

The tutorial describes the concept of edit distances applied to research andcommercial contexts. We use Translation Edit Rate (TER), Levenshtein,Damerau-Levenshtein, Longest Common Subsequence and $n$-gram distances todemonstrate the frailty of statistical metrics when comparing text sequences.Our discussion disassembles them into their essential components. We discussthe centrality of four editing actions: insert, delete, replace and move words,and show their implementations in openly available packages and toolkits. Theapplication of edit distances in downstream tasks often assumes that theseaccurately represent work done by post-editors and real errors that need to becorrected in MT output. We discuss how imperfect edit distances are incapturing the details of this error correction work and the implications forresearchers and for commercial applications, of these uses of edit distances.In terms of commercial applications, we discuss their integration incomputer-assisted translation tools and how the perception of the connectionbetween edit distances and post-editor effort affects the definition oftranslator rates.

 

Quick Read (beta)

loading the full paper ...