Logo Lanfrica

Developing trustworthy language models with consistent predictions

Domain:

natural language processing

Record type:

paper
Creator:
Jan
Editor:
LukCue
Publisher:
University of Oxford
Host:avatar
Transformer-based Pre-trained Language Models (PLMs), which are trained with a vast amount of natural language text, have significantly propelled the rapid progress in the field of Natural Language Processing (NLP). These models have been effectively employed across diverse downstream tasks in various ways, such as fine-tuning or few-shot learning. Notably, they have demonstrated promising performance, even surpassing human capabilities in several downstream tasks. Derived from these outstanding experimental results, a prevailing belief has emerged asserting that PLMs contain extensive knowledge regarding the world and possess the capability to comprehend natural language. However, a number of investigations have delineated the inherently imperfect capability of PLMs for comprehending natural language. Through experiments across diverse spectrums, it was revealed that PLMs exhibit deficiencies in understanding negation expressions and number-related knowledge. An additional avenue of research ascertained that PLMs exhibit logically incorrect behaviours, which greatly diverge from those demonstrated by humans. These fallacious behaviours have presented a substantial concern, as they undermine the trustworthiness of PLMs, potentially impeding their applications across industries, particularly in risk-sensitive areas, such as medical, financial, and legal domains. In this context, this dissertation centres its attention on enhancing the trustworthiness of Language Models (LMs) in terms of the consistency perspective. Although there have been several attempts that have explored the consistent behaviours of LMs and endeavoured to construct enhanced models equipped with heightened consistency, these studies have critical drawbacks. First, the definition of consistency has diverged across multiple studies, leading to fragmented investigations that hindered the achievement of unified and comprehensive evaluations. The diverse definitions have also resulted in incomprehensive alleviation methods that exclusively focus on certain types of consistency and cannot address other consistency types. Second, methodologies that have been proposed to improve consistent behaviours require substantial resource allocation. The most widely used techniques are data augmentation, which involves the collection of additional data that adheres to a designated consistency type, and consistency regularisation, where an additional training loss function is introduced to penalise undesirable behaviours through the utilisation of the augmented data. While the automatic collection of supplementary data to address certain consistency types is attainable, as exemplified by the utilisation of symmetry properties, the majority of these strategies mandate ample linguistic resources or significant human endeavours to ensure a high standard of data quality. Consequently, these mitigation methods become comparatively less feasible for languages with limited resources or for researchers operating under resource constraints. Furthermore, the incorporation of the augmented data and auxiliary regularisation terms for updating the parameters of LMs presents a challenge in terms of computational resources required for training. This is particularly pronounced due to the substantial increase in the size of contemporary LMs, as evidenced by GPT-4. This dissertation aims to address the aforementioned drawbacks of previous studies regarding LMs' consistent behaviours. First, I present a comprehensive definition of LMs' consistency based on the notion of behavioural consistency and propose a systematic categorisation that classifies various consistency types addressed in prior studies into three exclusive categories. Second, I introduce a unified benchmark dataset, which is designed to facilitate a comprehensive evaluation encompassing multiple consistency types across diverse downstream tasks. Third, I propose a practically efficient and feasible approach for enhancing LMs' consistent behaviours. Specifically, the method facilitates LMs to capture the precise meaning of language by learning conceptual roles from dictionary data. Subsequently, the LM enhanced with conceptual role information is incorporated with existing LMs, aiming to fully employ acquired knowledge. This is achieved through parameter integration in a resource-efficient manner that requires minimal computational resources. The extensive experimental results indicate two important findings. Above all, the investigation of this dissertation ascertains that, regardless of their size, architecture design, and training objectives, modern LMs exhibited inconsistent behaviours in many test cases, which is distinguished by consistency types and downstream tasks. None of them demonstrated coherently consistent behaviours in every test case. Subsequent findings demonstrate that the proposed alleviation method concurrently enhances a multitude of consistency types, a capability that previous methods lack. Furthermore, the experimental results attest to the reduction in computational resources achieved by the proposed approach and its wide applicability to low-resource languages beyond English.