LLMLingua-2, Originally developed and implemented in Python by Microsoft, is a small-size yet powerful prompt compression method.
llmlingua-2-js, ported by atjsh, is a pure JavaScript/TypeScript implementation of LLMLingua-2, designed to run in web browsers and Node.js environments.
You can try it on the GitHub Pages without any installation.
The source code for demo is available in the examples directory. You can run it locally using the following command:
cd examples/react-vite-webgpu
yarn install
This implementation depends on the following libraries:
Especially, the @huggingface/transformers library utilizes various computational optimizations to achieve high performance. Please consult if the running environment supports the minimum requirements from these libraries.
You can use the library by downloading the library from npm and importing it into your JavaScript/TypeScript project.
First, install the dependencies:
npm install @huggingface/transformers js-tiktoken
Then, install the library:
npm install @atjsh/llmlingua-2
Pass an explicit transformerJSConfig to the factory. These portable
configurations favor predictable output across supported runtimes:
const nodeTransformerJSConfig = {
device: "cpu",
dtype: "fp32",
} as const;
const browserTransformerJSConfig = {
device: "wasm",
dtype: "fp32",
} as const;
Hardware acceleration is opt-in and should be tested on each target platform:
| Runtime | device |
dtype |
|---|---|---|
| Windows Node.js with DirectML | "dml" |
"fp32" |
| Linux Node.js with NVIDIA CUDA | "cuda" |
"fp32" |
| Compatible browsers or macOS Node.js | "webgpu" |
"fp32" |
The library passes the configuration to Transformers.js without detecting a backend or retrying with a fallback. For an XLM-RoBERTa model stored with one external data file, also pass:
modelSpecificOptions: {
use_external_data_format: { "model.onnx": 1 },
}
You can choose between models based on your needs.
| Model | Size | Pros | Cons | Public Model |
|---|---|---|---|---|
| TinyBERT | 57.1 MB | Very small, fast |
Lower accuracy than larger models |
atjsh/llmlingua-2-js-tinybert-meetingbank |
| MobileBERT | 99.2 MB | Small, optimized for mobile |
Moderate accuracy, tradeoff in depth |
atjsh/llmlingua-2-js-mobilebert-meetingbank |
| BERT | 710 MB | Faster, smaller size |
Lower accuracy than XLM-RoBERTa |
Arcoldd/llmlingua4j-bert-base-onnx |
| XLM-RoBERTa | 2240 MB | High accuracy | Slower, slightly larger in size |
atjsh/llmlingua-2-js-xlm-roberta-large-meetingbank |
Learn More about the performance of each model (actual performance may vary).
The model files will be downloaded automatically by default.
For more details on how to use the library, please refer to the API reference documentation.
Unit tests are not available at the moment.
E2E tests are partially available in following directories:
src/e2eexamples/**See LICENSE for details.
This software includes other software related under the following licenses: