You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
This repository was archived by the owner on Feb 11, 2026. It is now read-only.
Repository navigation
This repository was archived by the owner on Feb 11, 2026. It is now read-only.
I would like to include a service in our stack that will perform tokenization on a subset of models. I think this would be helpful for users to understand more about their context size constraints, and to help them when writing their QNA pairs.
After thinking about this, my implementation in #210 is totally insufficient. VLLM does not serve us anything out of the box for doing tokenization without doing the inference. This leaves us with 2 options for something like this:
Lazy Implementation: Do the inference and just extract: .usage.prompt_tokens.
This should work for most easy contexts. However it has 2 drawbacks.
First is resource utilization - it would doing an inference request. Not just that, the goal of this feature was to help identify context side to determine sizing of QNA pairs to make sure it fits within the context window, so you could very feasibly be hitting the limits on the size of an inference for a significant amount of requests .
The second is that since your doing chat inference, some of the tokens will be taken up by the response. This means that this significantly increases the amount of tokens taken up, which would push a significant amount of requests out of the context window. This is not to mention the gateway timeouts, which I hit with a 1700 hundred word essay I fed it context request.
Robust solution: Build our own server for this
The solution would container 3 components.
Build a simple server for tokenization at the point.
We would build a simple flask, uvicorn, or other server. We would expose this on a separate port, and it would either use the VLLM configuration to lookup all models or it would curl the local model server deployed on the same machine and grab the model options from /v1/models. From there it use tiktoken to tokenize from the model path and serve those results. This is dependent on us being able to deploying things to the backend service or migrating this to our controll.
Create a proxy mapping for this to our endpoints
As with the chat completions we would map the tokenization server to a proxy endpoint
Write some client typescript code to use this tokenization
This would be consumed in the QNA section of the knowledge submission process.
Other considerations
The productized models will ship with 128k context window but not sure this applies to the upstream. If that is not the case for the upstream models, we should highly consider something like this.
Where does this rank on our list of priorities @vishnoianil ?
From upstream perspective, We will only support granite model (assuming tokenizer is same across different version of the granite model), so i think a simple service, that exposes /tokenlen API that returns the number of token for granite model would be pretty lightweight and should satisfy our requirement.
At this point of time the priority of this task is not too high, so we can prioritize it for the upstream release-1.2.
I would like to include a service in our stack that will perform tokenization on a subset of models. I think this would be helpful for users to understand more about their context size constraints, and to help them when writing their QNA pairs.
cc @nerdalert @vishnoianil @jjasghar