jevchat samples next symbols from Jev API via choice/bisect/buckets/refine strategies
jevchat demonstrates token-by-token sampling against an external probability API, revealing both the generality of next-symbol prediction and its prohibitive cost when state is not cached. The four strategies and beam-search mode expose concrete trade-offs between API calls and output quality that standard integrated LLMs avoid.
The repository implements four query strategies against the Jev probability API. Choice sends one full-alphabet request per step. Bisect performs binary search down to groups of 20 then resolves locally. Buckets partitions large vocabularies across parallel questions with an OTHER escape. Refine first buckets then rescoring a nucleus of winners. Each mode trades API calls for distribution accuracy while exposing live per-step symbol probabilities and generation rates in symbols per second. Benchmarks in the repo show choice without shuffling requires the fewest calls yet suffers position bias. Ensemble averaging of four shuffles raises cost linearly. Beam search at width 3 eliminates temperature and nucleus sampling, ranking candidates by joint probability instead. Generation remains token-by-token because every symbol draw triggers a fresh API round-trip, producing visible latency and partial outputs that stay in history on interrupt. Standard autoregressive models cache hidden states across steps and emit full distributions in one forward pass. jevchat cannot. The resulting slow, sometimes incoherent output illustrates both the universality of next-symbol prediction and its practical limits when the oracle is external and uncached. In an emotional context the same mechanism that can echo a known voice also exposes how many separate probability queries are required to produce even short coherent replies. Operationally the experiment quantifies why production chat systems moved from per-token external oracles to integrated models with KV caching. Future Jev API revisions that expose batched or cached next-token logits would collapse the observed latency gap without changing the underlying sampling logic.
Jev API: batch next-token endpoint ships before 2025-Q3, cutting median latency below 50 ms per symbol at width 1.
Sources (3)
- [1]Primary Source(https://github.com/kyle-pena-nlp/jevchat/)
- [2]Supporting Source(https://arxiv.org/abs/1706.03762)
- [3]Supporting Source(https://arxiv.org/abs/1904.09728)