I assume that this is "$ spent on search + $ spent on LLM" < budget, but how do you handle the LLM spending more than you would expect on a request? Or is this handled by max_tokens and some form of pricing table? (and if so, how does caching play a role?)
I'm glad your numbers are honest! For a moment I thought, hey, maybe this person's numbers are lying to me... but it turned out they were not so thank you!
conception 3 hours ago [-]
They are honest because they load bear the seam.
recroad 11 hours ago [-]
That is a LOT of code for a pretty basic feature.
reindeer2 7 hours ago [-]
[dead]
Rendered at 08:26:29 GMT+0000 (Coordinated Universal Time) with Vercel.
I assume that this is "$ spent on search + $ spent on LLM" < budget, but how do you handle the LLM spending more than you would expect on a request? Or is this handled by max_tokens and some form of pricing table? (and if so, how does caching play a role?)
[0]: https://www.datamole.ai/
I'm glad your numbers are honest! For a moment I thought, hey, maybe this person's numbers are lying to me... but it turned out they were not so thank you!