Gremlin Introduces Disaster Recovery Testing for Enhanced Cloud Resilience
- Securityish
- Tools & Best Practices
Quick Summary
The Securityish Brief
Gremlin, known for its proactive reliability platform, has introduced Disaster Recovery Testing to help organizations effectively manage failovers across zones, regions, and datacenters. This product was developed in light of notable cloud outages in 2025, including the AWS us-east-1 zone outage that impacted 70,000 companies and resulted in an estimated $581 million in losses. The testing allows businesses to simulate major failures and verify that their backup mechanisms function correctly.
Fred Bull, Security Officer at Gremlin, emphasizes that the product enables companies to create failure conditions, akin to fuzz testing, to assess the resilience of their systems. This includes checking how authentication systems respond to various dependency failures, such as NTP failures or certificate expirations. Such proactive measures are crucial for organizations, especially as they prepare for public offerings.
Gremlin’s Disaster Recovery Testing offers several key features, including company-wide testing from a central command center, enhanced safety measures through automatic health checks, and detailed reliability reports that highlight service performance weaknesses. These reports are especially beneficial for companies preparing their S-1 filings for the SEC or their 10-K annual filings, as they provide a comprehensive overview of digital resilience efforts.
Gremlin has collaborated with numerous Fortune 1000 companies, including four of the top five U.S. banks, to refine their disaster recovery strategies. This collaboration has informed Gremlin’s approach, allowing them to offer tailored guidance to clients seeking to optimize their testing strategies.
Sreekanth Rajagopal from Visa Cross-Border Solutions notes the importance of continuous availability and performance of applications, especially during outages. The Disaster Recovery Testing product provides a centralized method for validating and demonstrating resilience against catastrophic events, ensuring that services remain online.
Key Takeaways
- Evaluate your organization’s business continuity strategy to ensure it accounts for potential cloud outages.
- Consider implementing disaster recovery testing to simulate failover scenarios and verify system resilience.
- Review your authentication systems to ensure they can handle dependency failures effectively.
- Utilize reliability reports to identify weaknesses in your service performance and prioritize remediation efforts.
- Stay informed about cloud service outages and adjust your disaster recovery plans accordingly.
Key Terms & Concepts
- Disaster Recovery Testing: In this article, Disaster Recovery Testing refers to a new product by Gremlin that allows organizations to test their failover capabilities during major outages.
- Business Continuity Plan: A Business Continuity Plan is a strategy that outlines how an organization will continue to operate during and after a disaster.
- Reliability Reports: Reliability Reports are detailed documents generated by Gremlin that assess service performance and identify areas for improvement.
Your 5-Minute Securityish Brief
A weekly digest of cybersecurity news, phishing alerts, privacy tips, and emerging threats, simplified so anyone can understand what matters and why.
Securityish
Securityish explains cybersecurity, scams, data breaches, and privacy risks in simple language so you know what’s happening and how to protect yourself.
Navigation
Your 5-Minute Cybersecurity Brief
A weekly digest of cybersecurity news, phishing alerts, privacy tips, and emerging threats, simplified so anyone can understand what matters and why.