Text Ranking
sentence-transformers
Safetensors
Transformers
multilingual
t5gemma2
text2text-generation
reranker
encoder-decoder
FBNL
matryoshka
retrieval
RAG
cosyy commited on
Commit
f25c6a9
·
verified ·
1 Parent(s): 03ea0fd

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +107 -0
README.md CHANGED
@@ -339,6 +339,113 @@ rankings: [{'corpus_id': 0, 'score': 0.9867772459983826}, {'corpus_id': 1, 'scor
339
 
340
  ```
341
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
342
  #### Ablation on multi-stage training
343
 
344
  Across all three model sizes and all seven compression ratios, performance on BEIR and MIRACL improves consistently from Stage 1 to Stage 3, demonstrating the effectiveness of our multi-stage training pipeline. Concretely, Stage 1 establishes a robust foundation for document reranking, distillation in Stage 2 substantially improves performance, and Stage 3 yields further modest gains. More importantly, robustness to compression generally improves across the three training stages. For example, from Stage 1 to Stage 3, the performance retention of KaLM-Reranker-V1-Nano at r = 128 relative to r = 2 increases from 92.88% to 93.80% on BEIR and from 90.93% to 92.15% on MIRACL.
 
339
 
340
  ```
341
 
342
+ ### Using vLLM
343
+ An experimental single-GPU adapter is available for offline
344
+ `LLM.classify()` reranking and optional FastAPI serving. It reuses the original
345
+ checkpoint without adding or modifying model weights.
346
+
347
+ The adapter has been validated with Python 3.12, vLLM 0.19.1, Transformers
348
+ 5.6.2 and CUDA BF16:
349
+
350
+ ```bash
351
+ conda create -n kalm-vllm python=3.12 -y
352
+ conda activate kalm-vllm
353
+ pip install "vllm==0.19.1" "transformers==5.6.2"
354
+
355
+ hf download KaLM-Embedding/KaLM-Reranker-V1-Large-R2 \
356
+ --local-dir ./KaLM-Reranker-V1-Large-R2
357
+ pip install ./KaLM-Reranker-V1-Large-R2/vllm_support --no-deps
358
+ export VLLM_PLUGINS=kalm_t5gemma2
359
+ ```
360
+
361
+ If you modify the `vllm_support` source code for some reason
362
+ (e.g. changing the model path in `vllm_support/src/kalm_t5gemma2_vllm_plugin/constants.py`),
363
+ please run `pip install ./KaLM-Reranker-V1-Large-R2/vllm_support --no-deps --force-reinstall` to refresh the patch.
364
+
365
+ Offline Python:
366
+
367
+ ```python
368
+ from kalm_t5gemma2_vllm_plugin import KaLMVLLMReranker
369
+
370
+ query = "What is the capital of China?"
371
+ documents = [
372
+ "The capital of China is Beijing.",
373
+ "Gravity attracts bodies toward one another.",
374
+ ]
375
+
376
+ with KaLMVLLMReranker(
377
+ "KaLM-Embedding/KaLM-Reranker-V1-Large-R2",
378
+ query_max_length=512,
379
+ document_max_length=1024,
380
+ encoder_chunk_size=4,
381
+ ) as reranker:
382
+ print(reranker.rank(query, documents))
383
+ ```
384
+
385
+ Offline CLI:
386
+
387
+ ```bash
388
+ kalm-vllm-rerank --return-margin
389
+ ```
390
+
391
+ To deploy the online service, install the HTTP dependencies and keep the
392
+ server running in the first terminal:
393
+
394
+ ```bash
395
+ pip install "fastapi>=0.136,<0.137" "uvicorn>=0.46,<0.47"
396
+ export CUDA_VISIBLE_DEVICES=0
397
+ export VLLM_PLUGINS=kalm_t5gemma2
398
+
399
+ kalm-vllm-serve \
400
+ --host 0.0.0.0 \
401
+ --port 8000 \
402
+ --model KaLM-Embedding/KaLM-Reranker-V1-Large-R2 \
403
+ --query-max-length 512 \
404
+ --document-max-length 1024 \
405
+ --encoder-chunk-size 4 \
406
+ --max-model-len 2048
407
+ ```
408
+
409
+ In a second terminal, check the server:
410
+
411
+ ```bash
412
+ conda activate kalm-vllm
413
+ kalm-vllm-client --base-url http://127.0.0.1:8000 --health
414
+ ```
415
+
416
+ Use `/rerank` for one query and a list of documents. Results are sorted by
417
+ score:
418
+
419
+ ```bash
420
+ kalm-vllm-client \
421
+ --base-url http://127.0.0.1:8000 \
422
+ --endpoint rerank \
423
+ --json-file ./KaLM-Reranker-V1-Large-R2/vllm_support/examples/rerank_request.json \
424
+ --return-margin \
425
+ --top-k 10
426
+ ```
427
+
428
+ Use `/score` to score a batch of independent query-document pairs. Results
429
+ preserve the input order and optional IDs:
430
+
431
+ ```bash
432
+ kalm-vllm-client \
433
+ --base-url http://127.0.0.1:8000 \
434
+ --endpoint score \
435
+ --json-file ./KaLM-Reranker-V1-Large-R2/vllm_support/examples/score_request.json \
436
+ --return-margin
437
+ ```
438
+
439
+ The default output is `P(yes)`. Set `return_margin=true` to also receive
440
+ `yes_logit - no_logit`; the client flag `--return-margin` applies the same
441
+ setting to a JSON file request. The supported encoder chunk sizes are
442
+ `1, 2, 4, 8, 16, 32`, with `4` as the default.
443
+
444
+ This adapter uses vLLM's plugin, scheduling and pooling interfaces while the
445
+ T5Gemma2 semantic forward still runs through Transformers. It is not vLLM's
446
+ native HTTP `/score` implementation or a complete vLLM-native kernel port.
447
+ See [the complete installation, API and troubleshooting guide](./vllm_support/README.md).
448
+
449
  #### Ablation on multi-stage training
450
 
451
  Across all three model sizes and all seven compression ratios, performance on BEIR and MIRACL improves consistently from Stage 1 to Stage 3, demonstrating the effectiveness of our multi-stage training pipeline. Concretely, Stage 1 establishes a robust foundation for document reranking, distillation in Stage 2 substantially improves performance, and Stage 3 yields further modest gains. More importantly, robustness to compression generally improves across the three training stages. For example, from Stage 1 to Stage 3, the performance retention of KaLM-Reranker-V1-Nano at r = 128 relative to r = 2 increases from 92.88% to 93.80% on BEIR and from 90.93% to 92.15% on MIRACL.