CurrentWave AI - Intermittent slow chat responses – Incident details

Intermittent slow chat responses

Resolved
Operational
Started 9 days agoLasted 4 days

Affected

CXplain.ai

Degraded performance from 2:31 PM to 2:16 PM, Operational from 2:16 PM to 3:34 PM

Updates
  • Resolved
    UTC
    Resolved

    This incident has been resolved. The affected database cluster has been healthy for 72 hours, and our platform vendor confirms that all systems are stable. Thank you for your patience as we worked through this issue, and we sincerely apologize for the inconvenience.

  • Update
    UTC
    Update

    All CWAI systems continue to be running normally. Due to the severity of the database cluster issue on 9/10, we are continuing to monitor the health of the system closely alongside our platform vendor's product team, to ensure the system remains stable. We'll continue to provide updates here until we are certain that the issue is completely mitigated.

  • Update
    UTC
    Update

    All CWAI systems are up and running normally following yesterday's intermittent slow or broken chat responses caused by a failed hardware node in one of our database clusters. The team is currently closely monitoring system and database performance this morning, but all systems are looking back to normal.

  • Monitoring
    UTC
    Monitoring

    The database cluster has been repaired and all services are back to normal. We will be closely monitoring the performance of the database and system through the evening, and into the morning. Thank you for your patience, and we apologize again for the inconvenience.

  • Update
    UTC
    Update

    We have an update from our platform vendor that they need to complete a full rebuild of the damaged database cluster.  Unfortunately this rebuild requires the database to temporarily be offline for additional time.  The latest ETA from the vendor is that the cluster rebuild should be completed by 7:30pm Pacific time.  We continue to be very sorry for this inconvenience, and will continue to provide updates as we get them.

  • Update
    UTC
    Update

    Our platform provider is currently implementing a fix to the damaged database cluster.  The vendor is estimating approximately 2 hours for the fix to be completed, with an ETA of 4:30pm Pacific time.  We are extremely sorry for the inconvenience, and will post updates here as soon as they become available.

  • Update
    UTC
    Update

    Per our platform vendor: multiple teams are actively working on mitigating the issue and are working to ensure the issue is resolved as quickly as possible.  We very much apologize for the inconvenience.  We'll continue to post updates as we get them.

  • Update
    UTC
    Update

    The CWAI system continues to be up and running, but with degraded chat (question and answer service) caused by a failed hardware node in one of our database clusters.  The system is up, but is degraded.  Our platform provider is aware of the issue, has identified the bad component, and is actively working on a fix.  Most users are able to use the system normally, but with some answers being very slow or some answers failing completely.  The rest of the CWAI system is up and running.  We do not yet have an ETA from the platform provider for the mitigation to be complete, but we'll post updates here as they become available.

  • Update
    UTC
    Update

    The platform vendor has acknowledged an underlying database hardware issue affecting our service. They are working on a fix. We do not have an ETA yet for full recovery. We will continue to provide updates as we get more information.

  • Identified
    UTC
    Identified

    We are aware of issues where users are intermittently receiving slow responses to chat requests, and periodic no responses. We have isolated to a particular database component and are working with the platform vendor to investigate. We will continue to provide updates here as we get more information.