
In a new InfoQ presentation, American Express executive Matthew Liste distilled more than 20 years of mission-critical infrastructure experience into 12 principles for building resilient platforms. Speaking from a career spanning Goldman Sachs, JPMorganChase and American Express, Liste argued that platform teams succeed when they hide complexity, stay evergreen, and treat reliability, security and scalability as non-negotiable foundations.
What Matthew Liste means by a platform
Liste began by defining a platform as a set of integrated technologies that forms a foundation for applications built on top. In his view, cloud services such as AWS, Azure and GCP are among the clearest examples, but the idea applies anywhere software depends on shared infrastructure.
He also stressed that platforms are usually invisible when they work well. Using plumbing as his analogy, he said consumers take infrastructure for granted until it breaks. That, he argued, is the hallmark of strong platform design: “no one knows you’re there” because the experience is intuitive and the complexity is safely hidden.
Why the hidden work matters
Liste’s point was not simply that infrastructure is complicated. It is that platform teams are paid to absorb that complexity so developers can focus on business problems instead of operational detail. As he put it, most teams specialize because someone else has already built the platform layers they rely on.
- Platforms should feel obvious and transparent to the consumer.
- Complexity should be managed underneath the surface, not pushed onto developers.
- Good platforms become invisible in everyday use until something goes wrong.
The 12 principles behind resilient infrastructure platforms
Liste organized the talk around 12 principles he said he has used repeatedly to build infrastructure for large enterprises. He noted that the list is not ranked in order of importance, but grouped to create a practical framework for senior architects and engineering leaders.
Among the core ideas: deliver an intuitive experience, build common and interchangeable components, and use the “three S’s” of stability, security and scalability. He also emphasized evergreen maintenance, avoiding undifferentiated heavy lifting, and being opinionated about what a platform should and should not do.
Stability, security and scalability come first
For mission-critical systems, Liste said the “three S’s” are non-negotiable. A platform that fails frequently will not be trusted; one that is insecure creates unacceptable risk; and one that cannot scale will eventually undermine the first two.
He tied that directly to the financial-services environment, where American Express credit card authorization systems run at six nines or seven nines of availability, according to his talk. Other systems, such as mobile apps, may tolerate more downtime, but the expectation rises sharply as business criticality increases.
He also warned that many platforms seem successful until usage grows unexpectedly. At that point, bottlenecks appear, and scaling issues turn into reliability issues. His message was blunt: if people like the product, they will use it more, so teams need to design for that outcome from the start.
Staying evergreen is harder than it sounds
One of Liste’s strongest themes was the need to keep platforms current. He described “be evergreen” as one of the hardest disciplines to maintain, because teams are always tempted to delay upgrades in favor of visible feature work.
That delay, he said, eventually creates more client impact, not less. When a platform falls behind version by version, upgrades become disruptive and expensive. His preferred approach is to keep change continuous, predictable and as invisible to clients as possible.
He also introduced a shorthand for platform currency management: 0114. In his framing, a healthy platform should require zero people manually maintaining currency, should be upgradable across the fleet in less than a day, and should be cycled at least every 14 days.
- 0 people manually maintaining currency
- 1 day or less to upgrade the fleet
- 14 days maximum between cycles
Opinionated platforms reduce waste
Liste repeatedly returned to the idea that platform teams must say no. In his view, enterprise IT often struggles with client listening compared with customer-facing product companies, but listening does not mean building everything requested. Instead, platform owners need to deliver what provides the most value to the most users.
That means retiring technical debt, shutting down old features, and resisting the urge to support every edge case indefinitely. He argued that platform teams should be experts in the platform itself, then make a decisive call about which features belong in the shared service and which do not.
He also described the long lifespan of infrastructure decisions. Once a platform attracts users, it can be difficult to remove, even if the team later regrets building it. His advice: think carefully before creating something that will become a multi-year commitment.
How platform teams should iterate
Although Liste advocated long-term thinking, he also made a case for fast iteration during development. He said platform teams should “fail quick, fail often” because the best designs usually emerge through repeated learning rather than early certainty.
He illustrated that point with the evolution of container orchestration. His teams adopted Linux container primitives early, then built orchestration themselves before moving to Mesosphere and later Kubernetes. That sequence, he said, helped them gain experience ahead of broader industry adoption.
The lesson was not to keep migrating for its own sake, but to learn fast while preserving client fidelity. In his telling, the underlying container primitives remained consistent, which made the transitions manageable for consumers.
Shared responsibility and clear boundaries
Another major theme was shared responsibility. Liste said platform owners must be explicit about what they provide and what customers must handle themselves, ideally through contracts, APIs, SLOs and SLAs that are easy to understand.
He pointed to cloud providers as examples of clear shared-responsibility models, noting that well-defined boundaries help prevent surprises during outages. He also argued that platform providers can only offer fixed shapes, not artisanal one-off solutions for every consumer.
Abstract, but do not hide the system
Liste’s principle of “abstract, don’t obfuscate” focused on giving users multiple ways to interact with a platform, whether through APIs, SDKs, Terraform or a UI. The key, he said, is to expose enough detail for troubleshooting and customization rather than concealing how the system works.
He said that transparency matters for two reasons: first, when something breaks, clients need to know what happened under the hood; second, users sometimes need to modify generated infrastructure definitions instead of being trapped by a black box.
Open source and culture as multipliers
Two of Liste’s most emphatic points were about open source and culture. He said platform builders rely heavily on open-source software and open standards because they allow teams to focus on integration and differentiation rather than reinventing every layer.
Finally, he called culture the principle that makes the rest possible. Great culture, he said, builds great teams, and great teams build great products. That includes empowering teams to make decisions, encouraging diversity of thought, and managing team composition deliberately rather than assuming groups will self-organize optimally.
For Liste, culture is where resilient platforms begin: not with tooling alone, but with teams that can work together over years, absorb complexity, and keep systems dependable for the people who rely on them every day.
Source: Original report
Was this helpful?
Explore more: DevOps Services More Cloud & DevOps Tech News
Last Modified: October 7, 2026 at 10:33 pm
0 views
