Serverless pain is usually a systems problem
A Lambda timeout may begin in a partitioning choice. A high Athena bill may begin in an event schema. A fragile deployment may come from ownership boundaries rather than the cloud service itself. Modernization works when the whole path is visible enough to change one part without surprising the rest.
Evidence before migration
My production background includes event and analytics systems operating at hundreds of millions of events per tenant, including query work that improved performance by 12x. Those numbers matter here as evidence of the kind of constraint I have worked inside—not as a promise that every system will produce the same result.
A staged way forward
Map the expensive and failure-prone path with actual usage evidence.
Separate quick operational wins from changes that alter system boundaries.
Ship the smallest safe modernization slice and observe it under representative load.
Leave the team with clearer ownership, instrumentation, and a realistic next phase.
For a wider platform review, see Cloud Architecture and Optimization and the AppNavi case study.