Cargando...
Fecha
2025-06
Derechos de acceso
info:eu-repo/semantics/openAccess
Título de la revista
ISSN de la revista
Título del volumen
Editorial
Resumen
¿Tiene sentido que el espacio de significados representados por embeddings de palabras sea el mismo que el embeddings de grupos de palabras (sentencias en particular) obtenidos por composición de sus elementos? El marco de la Information Theory-based Compositional Distributional Semantics (ICDS) nos define cómo debe ser un espacio semántico vectorial de representación bajo el paradigma distribucional. Con un espacio de embeddings estáticos que cumple las propiedades exigidas por la teoría, ésta define cómo componer dos de estos vectores para obtener otro semánticamente representativo de tal composición. En este contexto de mínima composición es razonable obtener resultados semánticamente válidos. Pero, ¿hasta dónde podemos avanzar en una composición sucesiva de palabras manteniendo la coherencia semántica? O, en otros términos, ¿hasta qué punto tiene sentido representar con un único vector el significado de una frase arbitrariamente compleja en un espacio de representación léxico-semántico?
Nuestro trabajo parte de la idea de representar una sentencia como un conjunto de vectores en el espacio semántico distribucional, considerando que no existe un único sentido que pueda representar una frase de complejidad arbitraria en un único punto de ese espacio. Para ello, proponemos la aplicación de grafos de conocimiento mínimos (Short, Simple, Semantic, Knowledge Graphs, S3KG), asumiendo que cada sentencia es un conjunto de hechos (facts), expresables en general con un sujeto, una relación y un objeto o atributo, todos ellos con un mínimo de términos y cuya elementalidad podría hacer representativa en el espacio léxico-semántico la composición de los embeddings de los elementos que componen cada hecho, proyectándolo así en un punto en el espacio. Sería el conjunto de estos puntos/hechos el que conformaría la representación de la sentencia.
La interpretación del texto bajo un grafo nos aporta ventajas, como el incremento en la información representable; o la posible ponderación de cada tripleta según su aportación semántica al significado global de la sentencia; o una estructura sencilla de tres concisos elementos que minimiza el número de operaciones de composición y facilita explorar qué orden o qué ponderación, de existir, serían adecuados para la misma. Pero también presenta desventajas, como la mayor complejidad de una función de similitud adecuada, o establecer la cantidad de información representada, funciones estas que definirán el espacio de representación semántico en el que representemos las sentencias. Para la validación de nuestro modelo hemos realizado varios experimentos cuyos resultados ratifican en buena medida la estrategia de representación por hechos de las sentencias, a la vez que aportan señales sobre la conveniencia de ponderar diferencialmente, en la composición, los elementos (sujeto, relación y objeto/atributo) de las tripletas representativas de los hechos.
Does it make sense for the meaning space represented by word embeddings to be the same as that of embeddings for word groups (particularly sentences) obtained through the composition of their elements? The framework of Information Theory-based Compositional Distributional Semantics (ICDS) defines how a vector-based semantic representation space should behave under the distributional paradigm. Given a static embedding space that meets the theoretical requirements, ICDS specifies how to compose two such vectors to produce another that is semantically representative of that composition. In this minimal composition context, it is reasonable to expect semantically valid results. But how far can we go with successive word composition while maintaining semantic coherence? Or, in other words, to what extent does it make sense to represent the meaning of an arbitrarily complex phrase with a single vector in a lexical-semantic representation space? Our work is based on the idea of representing a sentence as a set of vectors in the distributional semantic space, considering that no single meaning can fully represent a phrase of arbitrary complexity in a single point of that space. To achieve this, we propose the use of minimal knowledge graphs (Short, Simple, Semantic Knowledge Graphs, S3KGs), assuming that each sentence is a set of facts —generally expressible with a subject, a relation, and an object or attribute, all with minimal terms— whose simplicity could make the composition of the embeddings of their elements representative in the lexical-semantic space, projecting each fact as a point in that space. The collection of these points/facts would then form the representation of the sentence. Interpreting text as a graph offers advantages, such as increased representable information, the possible weighting of each triplet based on its semantic contribution to the overall meaning of the sentence, or a simple three-element structure that minimizes the number of composition operations and facilitates exploring which order or weighting (if any) would be appropriate. However, it also presents disadvantages, such as the greater complexity of designing a suitable similarity function or determining the amount or content of information represented, functions that will define the semantic representation space in which sentences are modeled. To validate our model, we conducted several experiments whose results largely support the fact-based representation strategy for sentences while also providing insights into the benefits of differentially weighting, during composition, the elements (subject, relation, and object/attribute) of the triplets representing facts.
Does it make sense for the meaning space represented by word embeddings to be the same as that of embeddings for word groups (particularly sentences) obtained through the composition of their elements? The framework of Information Theory-based Compositional Distributional Semantics (ICDS) defines how a vector-based semantic representation space should behave under the distributional paradigm. Given a static embedding space that meets the theoretical requirements, ICDS specifies how to compose two such vectors to produce another that is semantically representative of that composition. In this minimal composition context, it is reasonable to expect semantically valid results. But how far can we go with successive word composition while maintaining semantic coherence? Or, in other words, to what extent does it make sense to represent the meaning of an arbitrarily complex phrase with a single vector in a lexical-semantic representation space? Our work is based on the idea of representing a sentence as a set of vectors in the distributional semantic space, considering that no single meaning can fully represent a phrase of arbitrary complexity in a single point of that space. To achieve this, we propose the use of minimal knowledge graphs (Short, Simple, Semantic Knowledge Graphs, S3KGs), assuming that each sentence is a set of facts —generally expressible with a subject, a relation, and an object or attribute, all with minimal terms— whose simplicity could make the composition of the embeddings of their elements representative in the lexical-semantic space, projecting each fact as a point in that space. The collection of these points/facts would then form the representation of the sentence. Interpreting text as a graph offers advantages, such as increased representable information, the possible weighting of each triplet based on its semantic contribution to the overall meaning of the sentence, or a simple three-element structure that minimizes the number of composition operations and facilitates exploring which order or weighting (if any) would be appropriate. However, it also presents disadvantages, such as the greater complexity of designing a suitable similarity function or determining the amount or content of information represented, functions that will define the semantic representation space in which sentences are modeled. To validate our model, we conducted several experiments whose results largely support the fact-based representation strategy for sentences while also providing insights into the benefits of differentially weighting, during composition, the elements (subject, relation, and object/attribute) of the triplets representing facts.
Descripción
Categorías UNESCO
Palabras clave
Citación
Isasa Cuartero, Agustín. Trabajo fin de Máster: "Representación Multi-Vectorial de Sentencias en el Espacio Léxico-Semántico: un modelo basado en Hechos a partir de Grafos de Conocimiento elementales" Universidad Nacional de Educación a Distancia (UNED), 2025
Centro
E.T.S. de Ingeniería Informática
Departamento
Lenguajes y Sistemas Informáticos

