Wherobots - API errors and workloads failing to start in Wherobots Cloud – Incident details

All systems operational

API errors and workloads failing to start in Wherobots Cloud

Resolved
Degraded performance
Started 21 days agoLasted about 2 hours

Affected

Wherobots Website

Degraded performance from 9:20 PM to 11:39 PM

Wherobots Catalog

Degraded performance from 9:20 PM to 11:39 PM

Wherobots Notebooks

Degraded performance from 9:20 PM to 11:39 PM

Wherobots Jobs

Degraded performance from 9:20 PM to 11:39 PM

Wherobots SQL API

Degraded performance from 9:20 PM to 11:39 PM

Updates
  • Resolved
    UTC
    Resolved

    This incident has been resolved. The identity-service degradation at our cloud provider ended at approximately 3:20 PM PDT (22:20 UTC), and API error rates, latency, and workload launches have been normal for over an hour since. Any SQL sessions, notebooks, or jobs that failed to start between 2:20 and 3:20 PM PDT can be retried normally. We apologize for the disruption and will follow up with the changes we're making to reduce our sensitivity to this class of provider issue.

  • Monitoring
    UTC
    Monitoring

    Our cloud provider's identity service has recovered. API error rates, latency, and workload launches have been normal since approximately 3:20 PM PDT (22:20 UTC), and all API instances are healthy. We are continuing to monitor before marking this incident resolved. Any workload launch that failed during the incident can be retried now.

  • Identified
    UTC
    Identified

    Since approximately 2:20 PM PDT (21:20 UTC) the Wherobots Cloud API has been experiencing intermittent errors, and some workloads are failing to start. Customers may see API requests and console actions fail or time out, and new SQL sessions, notebooks, and jobs may fail to launch. Tools that authenticate through the API (for example the MCP server) may also fail intermittently.

    We have identified the cause as a degradation in an identity service operated by our cloud provider. This is not caused by a change on our side, and no customer data is affected. Impact peaked between 2:20 and 2:45 PM PDT and has been intermittent since; some requests continue to fail while the provider recovers.

    If a workload fails to start, please retry — launches succeed once credentials are issued. We are monitoring closely and will post another update within 30 minutes or sooner if the situation changes.