# Remote fetch latency from LFS

**URL:** https://community.mindstudio.ai/t/remote-fetch-latency-from-lfs/549
**Category:** Bug Reports
**Created:** [April 22, 2025, 9:35am UTC](https://community.mindstudio.ai/t/remote-fetch-latency-from-lfs/549 "2025-04-22T09:35:39Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![heyhaiden](https://avatars.discourse-cdn.com/v4/letter/h/cdc98d/32.png) [@heyhaiden](https://community.mindstudio.ai/u/heyhaiden)
#### Post date: [April 22, 2025, 9:35am UTC](https://community.mindstudio.ai/t/remote-fetch-latency-from-lfs/549/1 "2025-04-22T09:35:39Z")

</div>

1. What happens:  
Two identical agent runs on the same event page using the same model (`llama-3.1-8b-instant-groq`) returned the same output, but with a massive performance difference.

Run ID: **d00e8da7** took ~35 seconds to return the first token  
Run ID: **a0386e03** completed in ~1.2 seconds.

1. Intended behavior:  
Consistent response latency across equivalent inputs. Remote file fetch or model queueing shouldn’t introduce ~30 sec variability.

2. Agent ID:  
d0f9126b-d4bb-4356-8283-50ae866c9ee6

3. Attachment:  
📎 See screenshot attached.

---

<div class="post-metadata">

### Author: ![Simon](https://yyz2.discourse-cdn.com/flex008/user_avatar/community.mindstudio.ai/simon/32/29_2.png) [@Simon](https://community.mindstudio.ai/u/Simon)
#### Post date: [April 22, 2025, 10:10am UTC](https://community.mindstudio.ai/t/remote-fetch-latency-from-lfs/549/2 "2025-04-22T10:10:10Z")

</div>

Thanks for getting in touch. Please may you also send this through to support@mindstudio.ai so we can take a deeper look into this.

Thank you!

---

<div class="post-metadata">

### Author: ![heyhaiden](https://avatars.discourse-cdn.com/v4/letter/h/cdc98d/32.png) [@heyhaiden](https://community.mindstudio.ai/u/heyhaiden)
#### Post date: [April 22, 2025, 11:02am UTC](https://community.mindstudio.ai/t/remote-fetch-latency-from-lfs/549/3 "2025-04-22T11:02:56Z")

</div>

Sent, thanks!

---

<div class="post-metadata">

### Author: ![sean](https://avatars.discourse-cdn.com/v4/letter/s/f0a364/32.png) [@sean](https://community.mindstudio.ai/u/sean)
#### Post date: [April 22, 2025, 11:23am UTC](https://community.mindstudio.ai/t/remote-fetch-latency-from-lfs/549/4 "2025-04-22T11:23:31Z")

</div>

Hi there, I took a deeper look and it appears this was a service interruption caused by Groq—it looks like they had a network issue from which it took a moment to recover. You can see in the attached logs screenshot (these are logs I pulled directly from Groq, not logs from MindStudio) that the first request fails entirely, then is automatically retried and takes a while to complete.

 ![image](https://canada1.discourse-cdn.com/flex008/uploads/mindstudio/original/1X/fde8aaf4030f2d0e4c0d77637be8e4c7e7cef019.png)

Unfortunately, while we try our best to deliver a consistent experience, the model providers are dealing with a lot of demand and we tend to see these sorts of issues from time to time (e.g., take a look at the “API” section on Anthropic’s status page to see how frequently they have outages: [https://status.anthropic.com/](https://status.anthropic.com/)).

Please let me know if that answers your query or if there is anything else I can do to help! Thanks!

---

<div class="post-metadata">

### Author: ![heyhaiden](https://avatars.discourse-cdn.com/v4/letter/h/cdc98d/32.png) [@heyhaiden](https://community.mindstudio.ai/u/heyhaiden)
#### Post date: [April 22, 2025, 11:56am UTC](https://community.mindstudio.ai/t/remote-fetch-latency-from-lfs/549/5 "2025-04-22T11:56:20Z")

</div>

Thanks Sean, totally understand. To clarify, when I see “Received first response token” in debugger, is this always referencing the model provider?

---

<div class="post-metadata">

### Author: ![sean](https://avatars.discourse-cdn.com/v4/letter/s/f0a364/32.png) [@sean](https://community.mindstudio.ai/u/sean)
#### Post date: [April 22, 2025, 12:24pm UTC](https://community.mindstudio.ai/t/remote-fetch-latency-from-lfs/549/6 "2025-04-22T12:24:22Z")

</div>

Correct! This is also seen in the Groq screenshot above as “TTFT” (time to first token). The latency column is then TTFT + how long it takes to generate the full result. Sometimes if TTFT is high it means there are availability or other network issues affecting the provider, where the model provider is struggling to get the request started. As opposed to low TTFT but high latency, which just means the model is taking a lot of time to do its work.

“Running model” = “MindStudio has sent the request to the model provider and the model provider has acknowledged that they have received the request”  
“Received first response token” = “We’ve started getting result data back from the model provider”  
“Received full result…” = “Model provider has finished responding and MindStudio is now processing the result”
