Repository navigation
Conversation
|
The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update. |
ArthurZucker
left a comment
There was a problem hiding this comment.
Nice!
Examples are cool, but imo a readme + pointing to the example space makes our changes minimal / better!
Let's try to update ptr_hash instead the maintainer is great and always responsive ! if possible !
| /// @param json - The contents of a `tokenizer.json` file. | ||
| /// @returns The tokenizer. | ||
| /// @throws If `json` is not valid JSON or does not describe a tokenizer. | ||
| pub fn from_json(json: &str) -> Result<Self, JsError> { |
There was a problem hiding this comment.
we have from_file and from_pretrained in other APIs
There was a problem hiding this comment.
Browsers don't expose the user's file system so from_file is a bit moot here
To load from the user's file system you would use <input type="file"> and a handle like that:
input.onchange = async () => {
const tokenizer = Tokenizer.from_json(await input.files[0].text());
};from_pretrained has been implemented, however
|
I've been compiling the 1.0 crates to wasm for my browser tokenizer playground and hit the same things, so a few notes from that:
On ptr_hash: I'm happy to open the upstream PR Arthur suggested. Since the timings only feed Update: PR created RagnarGrootKoerkamp/PtrHash#41 Update 2: merged and released in PtrHash 2.1.2 |
…embly Each model now loads and tokenises in its own Web Worker, sharing one compiled WebAssembly module, so typing no longer blocks on tokenising. Results for text that has changed since are dropped. Tokenizer files are fetched from the Hub in the worker and kept with the Cache API (the same cache transformers.js used). The vendored build comes from daaain/tokenizers@claude/wasm-integration (huggingface/tokenizers#2450 plus fixes in review upstream), and scripts/tokenizers-wasm/ rebuilds it and checks parity with transformers.js: all 14 tokenizers there match. Also: - A per-model "Clean up spaces before punctuation" switch, remembered in localStorage, showing the model's own default from tokenizer_config.json. It changes the text shown, never ids or count. - Error cards say whether a repo is missing or gated, or has no tokenizer.json. - The text box grows with field-sizing where supported, avoiding a full layout on every keystroke. - Removes transformers.js and the experiment it replaces. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RuhhXa22BzYGGVpkkVnoco
|
Thanks for porting the fix upstream in RagnarGrootKoerkamp/PtrHash#41 @daaain |
Shipping offset tracking is in the pipeline for v1, fyi |
No worries, it was a simple fix! BTW I got this branch's WASM version fully working in my online tokenizer and it's fast! Here's a readme about the bits I needed to do (and the script I used), with the PtrHash fix in and SIMD flag added, it's pretty much there (just the Gemma 2 fix and the decode_each method missing): https://github.com/daaain/online-llm-tokenizer/tree/main/scripts/tokenizers-wasm There's a transformers.js parity checker in there too if you're interested, but that might not be this branch's remit |
ArthurZucker
left a comment
There was a problem hiding this comment.
LGTM otherwise, trusting you entirely on the js / web but biding is alright! Le't sminimize files, code, etc if possilbe to have an as thin layer as possible!
There was a problem hiding this comment.
IMO if we can point to a space, I'd rather have that 👀 or an externa tokenizers example repo?
There was a problem hiding this comment.
I'm happy to volunteer my vanilla JS repo to be an example and to update to the canonical usage once this branch is merged and released!
| /// @param json - The contents of a `tokenizer.json` file. | ||
| /// @returns The tokenizer. | ||
| /// @throws If `json` is not valid JSON or does not describe a tokenizer. | ||
| pub fn from_json(json: &str) -> Result<Self, JsError> { |
TL;DR
A new bindings library using https://github.com/wasm-bindgen/wasm-bindgen to generate a npm package that can be used in a browser context
Goal would be to publish it as a npm package so users can simply
npm install tokenizers-web(or any other name) and use tokenize locally in their browser apps.TODO
Release package