Next Article in Journal
Identification of Black Spot Resistance in Broccoli (Brassica oleracea L. var. italica) Germplasm Resources
Next Article in Special Issue
REACT: Relation Extraction Method Based on Entity Attention Network and Cascade Binary Tagging Framework
Previous Article in Journal
The Effect of Inulin Addition on Rice Dough and Bread Characteristics
Previous Article in Special Issue
Causal Reinforcement Learning for Knowledge Graph Reasoning
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Prefix Data Augmentation for Contrastive Learning of Unsupervised Sentence Embedding

School of Mathematical Sciences, University of Electronic Science and Technology of China, Chengdu 611731, China
*
Author to whom correspondence should be addressed.
Appl. Sci. 2024, 14(7), 2880; https://doi.org/10.3390/app14072880
Submission received: 1 February 2024 / Revised: 24 March 2024 / Accepted: 25 March 2024 / Published: 29 March 2024

Abstract

This paper presents prefix data augmentation (Prd) as an innovative method for enhancing sentence embedding learning through unsupervised contrastive learning. The framework, dubbed PrdSimCSE, uses Prd to create both positive and negative sample pairs. By appending positive and negative prefixes to a sentence, the basis for contrastive learning is formed, outperforming the baseline unsupervised SimCSE. PrdSimCSE is positioned within a probabilistic framework that expands the semantic similarity event space and generates superior negative samples, contributing to more accurate semantic similarity estimations. The model’s efficacy is validated on standard semantic similarity tasks, showing a notable improvement over that of existing unsupervised models, specifically a 1.08% enhancement in performance on BERTbase. Through detailed experiments, the effectiveness of positive and negative prefixes in data augmentation and their impact on the learning model are explored, and the broader implications of prefix data augmentation are discussed for unsupervised sentence embedding learning.
Keywords: contrastive learning; sentence embedding; prefix data augmentation contrastive learning; sentence embedding; prefix data augmentation

Share and Cite

MDPI and ACS Style

Wang, C.; Lv, S. Prefix Data Augmentation for Contrastive Learning of Unsupervised Sentence Embedding. Appl. Sci. 2024, 14, 2880. https://doi.org/10.3390/app14072880

AMA Style

Wang C, Lv S. Prefix Data Augmentation for Contrastive Learning of Unsupervised Sentence Embedding. Applied Sciences. 2024; 14(7):2880. https://doi.org/10.3390/app14072880

Chicago/Turabian Style

Wang, Chunchun, and Shu Lv. 2024. "Prefix Data Augmentation for Contrastive Learning of Unsupervised Sentence Embedding" Applied Sciences 14, no. 7: 2880. https://doi.org/10.3390/app14072880

APA Style

Wang, C., & Lv, S. (2024). Prefix Data Augmentation for Contrastive Learning of Unsupervised Sentence Embedding. Applied Sciences, 14(7), 2880. https://doi.org/10.3390/app14072880

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop