@FabienPenchenat This is the answer that I received from the Forge LLM team:
Compute usage (GB-seconds) is a different cost element to the LLM usage (token based).
While LLM responses do take longer to respond, the running costs will therefore be twofold:
-
compute usage (longer running times)
-
LLM usage itself