提示词与指令现代嵌入模型E5、BGE、GTE、Qwen3-Embedding、Nomic、Instructor 等在编码时使用提示词/指令——如query: 、passage: 或Represent this sentence for retrieval:这样的短前缀——来表达任务意图。在 sentence-transformers 中提示词在训练和推理期间都有效库会自动保持它们对齐。提示词在分词之前被字面地前置到输入文本上。模型级提示词在模型上设置提示词使encode()、encode_query()、encode_document()自动使用它们modelSentenceTransformer(your-base-model)model.prompts{query:query: ,document:passage: ,}model.default_prompt_namedocument# 未指定时使用推理时q_embmodel.encode_query(What is the capital of France?)# 内部编码 query: What is the capital of France?d_embmodel.encode_document([Paris is the capital of France.,Berlin is the capital of Germany.])# 内部编码 passage: Paris... 等# 显式 prompt_nameembmodel.encode([some text],prompt_namequery)带提示词训练在训练参数上设置提示词使它们在训练期间自动应用于正确的列。四种形状越来越具体1. 单一提示词应用于所有地方argsSentenceTransformerTrainingArguments(...,promptsRepresent this sentence for similarity: ,)每个输入列都获得前缀。最简单通常对单任务训练就足够了。2. 按列提示词prompts{anchor:query: ,positive:passage: ,negative:passage: ,}键是列名。不同角色使用不同前缀。3. 按数据集提示词多数据集prompts{all-nli:Classify the entailment relationship: ,stsb:Score semantic similarity: ,}键是数据集名称与train_dataset字典键匹配。4. 按数据集 按列prompts{all-nli:,# all-nli无提示词msmarco:{# msmarco按列query:query: ,positive:passage: ,negative:passage: ,},}交叉编码器特有交叉编码器将提示词限制为单值或仅按数据集不支持按列。这是因为交叉编码器接受文本对而不是单独的列。argsCrossEncoderTrainingArguments(...,promptsRank this passage for the query: ,# 单一提示词)# 或argsCrossEncoderTrainingArguments(...,prompts{msmarco:...,gooaq:...},# 按数据集)稀疏编码器SPLADE稀疏编码器支持与双编码器相同的四种形状的提示词v5.x 中新增。相同的model.prompts {...}API。池化与提示词 token双编码器在使用带提示词前缀的 mean/last-token 池化时提示词 token 本身可以被包含在池化嵌入中默认行为。更简单大多数模型都这样做。被排除——只池化实际内容 token。要排除翻转 Pooling 模块上的include_prompt——通过isinstance找到它而不是固定索引因为多模态 / Router 流水线并不总是把 Pooling 放在位置 1fromsentence_transformers.sentence_transformer.modulesimportPoolingformoduleinmodel:ifisinstance(module,Pooling):module.include_promptFalse或者如果你的模型暴露了辅助方法使用它例如SparseEncoder.set_pooling_include_prompt(False)。何时排除当提示词纯粹是任务信号而你不希望它稀释语义表示时。大多数 E5 / BGE / GTE 模型保持提示词被包含。在使用prompts训练期间include_prompt设置会被尊重训练器还在内部跟踪include_prompt_lengths以便在include_promptFalse时池化正确跳过提示词 token。推理对齐如果你用prompts{query: query: , ...}训练你必须在模型上保存提示词在save_pretrained之前设置model.prompts args.prompts或让库通过模型卡数据自动完成。在推理时使用相同的提示词调用encode_query()/encode_document()保存的提示词会被应用。如果保存的模型的config_sentence_transformers.json有promptsdefault_prompt_name任何加载它的人都能免费获得正确的提示词。指令微调与提示词不同一些模型Instructor、Qwen3-Embedding使用指令——如Represent the biomedical query for retrieving relevant passages:这样更长的描述。在 sentence-transformers 中这些以相同方式建模只需将它们设置为提示词。模型卡的惯例是在字典中列出可用的指令/提示词model.prompts{msmarco-query:Represent the query for MS MARCO retrieval: ,msmarco-doc:Represent the passage for MS MARCO retrieval: ,sts:Represent the sentence for semantic similarity: ,classification:Represent the sentence for topic classification: ,}用户调用model.encode([...], prompt_namemsmarco-query)。陷阱用提示词训练但忘记在save_pretrained之前在模型上设置它们保存的模型不知道提示词。用户将不带前缀编码并得到糟糕的结果。修复保存前设置model.prompts args.prompts或使用带提示词信息的SentenceTransformerModelCardData(...)。在推理时使用encode_query()但没有用 “query” 提示词训练该方法只会调用常规encode不应用前缀这没问题。但文档暗示你使用它用户可能会困惑。尾随空格很重要query: 与query:——前者在真实文本前有一个尾随空格。检查你的训练数据以知道你的模型是用哪一个训练的。在多数据集训练中混合带提示词和不带提示词的数据是可以的只要你的prompts字典按数据集覆盖例如{dataset_a: query: , dataset_b: }。include_promptFalse配合小数据集模型可能欠拟合因为你实际上缩短了每个输入。通常保持为 True。