logo

Better prompt caching for GPT‑6

Posted by mehrdadrad |2 hours ago |2 comments

OutOfHere 15 minutes ago

The biggest continuing limitation I see is that the cached input has to be at least 1024 tokens. This is terrible. It means a lot of good prefixes that are smaller will go uncached for no good reason. The threshold should have been 128.