Publicación: A Novel Methodology for Enhancing Cross-language and Domain Adaptability in Temporal Expression Normalization
| dc.contributor.author | Sánchez de Castro Fernández, Alejandro | |
| dc.contributor.author | Araujo Serna, M. Lourdes | |
| dc.contributor.author | Martínez Romo, Juan | |
| dc.contributor.funder | Agencia Estatal de Investigación (España) | |
| dc.contributor.funder | Universidad Nacional de Educación a Distancia (UNED) | |
| dc.date.accessioned | 2026-01-27T08:50:00Z | |
| dc.date.available | 2026-01-27T08:50:00Z | |
| dc.date.issued | 2025-04-19 | |
| dc.description | The registered version of this article, first published in “Computational Linguistics 51 (4): 1303–1335", is available online at the publisher's website: https://doi.org/10.1162/COLI.a.12 | |
| dc.description | La versión registrada de este artículo, publicado por primera vez en “Computational Linguistics 51 (4): 1303–1335", está disponible en línea en el sitio web del editor: https://doi.org/10.1162/COLI.a.12 | |
| dc.description.abstract | Accurate temporal expression normalization, the process of assigning a numerical value to a temporal expression, is essential for tasks such as timeline creation and temporal reasoning. While rule-based normalization systems are limited in adaptability across different domains and languages, deep-learning solutions in this area have not been extensively explored. An additional challenge is the scarcity of manually annotated corpora with temporal annotations. To address the adaptability limitations of current systems, we propose a highly adaptable methodology that can be applied to multiple domains and languages. This can be achieved by leveraging a multilingual Pre-trained Language Model (PTLM) with a fill-mask architecture, using a Value Intermediate Representation (VIR) where the temporal expression value format is adjusted to the fill-mask representation. Our approach involves a two-phase training process. Initially, the model is trained with a novel masking policy on a large English biomedical corpus that is automatically annotated with normalized temporal expressions, along with a complementary hand-crafted temporal expressions corpus. This addresses the lack of manually annotated data and helps to achieve sufficient capacity for adaptation to diverse domains or languages. In the second phase, we show how the model can be tailored to different domains and languages using various techniques, showcasing the versatility of the proposed methodology. This approach significantly outperforms existing systems. | en |
| dc.description.provenance | Made available in DSpace on 2026-01-27T08:50:00Z (GMT). No. of bitstreams: 1 MartinezRomo_Juan_TemporalExpression_JUAN MARTÍNEZ ROMO.pdf: 1222597 bytes, checksum: aee350684f77feb1938cf6ad4abc9c99 (MD5) Previous issue date: 2025-04-19 | en |
| dc.description.sponsorship | This work has been funded by the following projects: OBSER-MENH (MCIN/AEI/10.13039/501100011033 and NextGenerationEU”/PRTR with identification TED2021-130398B-C21), SICAMESP (2023-VICE-0029), and by the project EDHER-MED (PID2022-136522OB-C21)”. | |
| dc.description.version | versión original | |
| dc.identifier.citation | Sánchez de Castro, A., Araujo, L., & Martinez-Romo, J. (2025). A Novel Methodology for Enhancing Cross-Language and Domain Adaptability in Temporal Expression Normalization. Computational Linguistics 51 (4): 1303–1335. https://doi.org/10.1162/COLI.a.12 | |
| dc.identifier.doi | https://doi.org/10.1162/COLI.a.12 | |
| dc.identifier.issn | 0891-2017 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14468/31578 | |
| dc.journal.issue | 4 | |
| dc.journal.title | Computational Linguistics | |
| dc.journal.volume | 51 | |
| dc.language.iso | en | |
| dc.page.final | 1335 | |
| dc.page.initial | 1303 | |
| dc.publisher | Massachusetts Institute of Technology Press | |
| dc.relation.center | E.T.S. de Ingeniería Informática | |
| dc.relation.department | Lenguajes y Sistemas Informáticos | |
| dc.relation.projectid | info:eu-repo/grantAgreement/AEIProyectos Estratégicos Orientados a la Transición Ecológica y a la Transición Digital 2021/TED2021-130398B-C21/ES/GELP: Generación mediante procesamiento del lenguaje de perfiles demográficos en redes sociales para la detección de riesgo de suicidio y su relación con otros problemas psicológicos | |
| dc.relation.projectid | info:eu-repo/grantAgreement/AEI/Proyectos de I+D+I (Generación de Conocimiento y Retos Investigación) 2022/PID2022-136522OB-C21/ES/Detección precoz de enfermedades de alto impacto mediante el procesamiento del lenguaje natural | |
| dc.rights | info:eu-repo/semantics/openAccess | |
| dc.rights.uri | http://creativecommons.org/licenses/by-nc-nd/4.0/deed.es | |
| dc.subject | 1203 Ciencia de los ordenadores | |
| dc.title | A Novel Methodology for Enhancing Cross-language and Domain Adaptability in Temporal Expression Normalization | en |
| dc.type | artículo | es |
| dc.type | journal article | en |
| dspace.entity.type | Publication | |
| relation.isAuthorOfPublication | 6c2a060d-18a1-4a12-b8d1-d41a6d335ec7 | |
| relation.isAuthorOfPublication | 77c4023e-4374-442a-9dfb-b9d4b609c31e | |
| relation.isAuthorOfPublication | 91b7e317-2a30-494f-98e9-3a0e026747b1 | |
| relation.isAuthorOfPublication.latestForDiscovery | 6c2a060d-18a1-4a12-b8d1-d41a6d335ec7 |
Archivos
Bloque original
1 - 1 de 1
Cargando...
- Nombre:
- MartinezRomo_Juan_TemporalExpression_JUAN MARTÍNEZ ROMO.pdf
- Tamaño:
- 1.17 MB
- Formato:
- Adobe Portable Document Format
Bloque de licencias
1 - 1 de 1
No hay miniatura disponible
- Nombre:
- license.txt
- Tamaño:
- 3.62 KB
- Formato:
- Item-specific license agreed to upon submission
- Descripción: