Cargando...
Fecha
2025-06-12
Derechos de acceso
info:eu-repo/semantics/openAccess
Título de la revista
ISSN de la revista
Título del volumen
Editorial
Resumen
Understanding and e!ectively utilizing sentence embeddings, a cornerstone of modern Natural Language Processing (NLP), is often hindered by their inherent black-box nature and the limitations of traditional aggregate evaluation metrics. This thesis addresses these challenges by presenting a novel Framework and Visual Analytics (VA) Tool for Exploring Compositionality in Sentence Embeddings. This tool empowers researchers to move beyond high-level performance scores and gain deeper, interactive insights into how various embedding models, composition functions, and similarity metrics influence textual representations.
The core contribution of this work is the VA tool itself, which integrates comprehensive visualizations with interactive filtering and detailed drill-down capabilities. To enable this granular analysis, I developed an experimental framework or pipeline that systematically processes and normalizes the outputs of diverse embedding models, ensuring consistent and comparable data for the VA application. This framework is a valuable artifact in its own right, facilitating the reproduction of results and the expansion of the tool’s dataset. Additionally, it motivates the following development of a VA tool.
The VA tool enhances embedding model interpretability by allowing visual exploration of embedding behavior across di!erent configurations and layers. It facilitates systematic comparison across models, even those with disparate architectures, within a unified analytical environment. Crucially, it moves beyond aggregate performance by focusing on the error gap between predicted and actual similarity scores. Through detailed examples, interactive error gap heatmaps, and an Alternative Functions Heatmap for specific challenging instances, the tool enables fine-grained evaluation, revealing nuanced model strengths and limitations often
obscured by summary statistics. This work provides an intuitive platform for diagnosing model failures, understanding representational biases, and fostering more informed decisions in the development and application of sentence embeddings.
Descripción
Categorías UNESCO
Palabras clave
Sentence Embeddings, Visual Analytics, NLP, Interpretability, Evaluation Metrics, Error Analysis, Experimental Framework, Contextual Embeddings
Citación
Delgado Camacho, David Eduardo. Trabajo de Fin de Máster: «A Framework and Visual Analytics Tool for Exploring Compositionality in Sentence Embeddings». Universidad Nacional de Educación a Distancia (UNED), 2025
Centro
E.T.S. de Ingeniería Informática
Departamento
Lenguajes y Sistemas Informáticos

