Skip to content
This repository was archived by the owner on Feb 11, 2026. It is now read-only.
This repository was archived by the owner on Feb 11, 2026. It is now read-only.

Tokenization service for supported models #209

Description

@Gregory-Pereira

I would like to include a service in our stack that will perform tokenization on a subset of models. I think this would be helpful for users to understand more about their context size constraints, and to help them when writing their QNA pairs.

cc @nerdalert @vishnoianil @jjasghar

Activity

  1. Gregory-Pereira commented on Dec 16, 2024

    @Gregory-Pereira
    ContributorAuthor

    After thinking about this, my implementation in #210 is totally insufficient. VLLM does not serve us anything out of the box for doing tokenization without doing the inference. This leaves us with 2 options for something like this:

    Lazy Implementation: Do the inference and just extract: .usage.prompt_tokens.

    This should work for most easy contexts. However it has 2 drawbacks.

    First is resource utilization - it would doing an inference request. Not just that, the goal of this feature was to help identify context side to determine sizing of QNA pairs to make sure it fits within the context window, so you could very feasibly be hitting the limits on the size of an inference for a significant amount of requests .

    The second is that since your doing chat inference, some of the tokens will be taken up by the response. This means that this significantly increases the amount of tokens taken up, which would push a significant amount of requests out of the context window. This is not to mention the gateway timeouts, which I hit with a 1700 hundred word essay I fed it context request.

    Robust solution: Build our own server for this

    The solution would container 3 components.

    Build a simple server for tokenization at the point.

    We would build a simple flask, uvicorn, or other server. We would expose this on a separate port, and it would either use the VLLM configuration to lookup all models or it would curl the local model server deployed on the same machine and grab the model options from /v1/models. From there it use tiktoken to tokenize from the model path and serve those results. This is dependent on us being able to deploying things to the backend service or migrating this to our controll.

    Create a proxy mapping for this to our endpoints

    As with the chat completions we would map the tokenization server to a proxy endpoint

    Write some client typescript code to use this tokenization

    This would be consumed in the QNA section of the knowledge submission process.

    Other considerations

    The productized models will ship with 128k context window but not sure this applies to the upstream. If that is not the case for the upstream models, we should highly consider something like this.

    Where does this rank on our list of priorities @vishnoianil ?

  2. vishnoianil commented on Dec 17, 2024

    @vishnoianil
    Contributor

    From upstream perspective, We will only support granite model (assuming tokenizer is same across different version of the granite model), so i think a simple service, that exposes /tokenlen API that returns the number of token for granite model would be pretty lightweight and should satisfy our requirement.

    At this point of time the priority of this task is not too high, so we can prioritize it for the upstream release-1.2.

  3. moved this to Backlog in UIon Feb 8, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions