| July 2026 | Major update. New chapters: Spark Agents (Upgrade Agent + Troubleshooting Agent), EMR Observability (CloudWatch, Prometheus/Grafana, logging architecture). Significantly rewritten chapters: Gathering Requirements (expanded with practical metric collection APIs, workload inventory, configuration mapping, and completion checklist); Amazon EMR on AWS Outposts (expanded with use-case guidance, architecture patterns, deployment considerations); Support for Your Migration (restructured with decision guide, expanded Partners section, modernized ProServe, added Community resources); Operational Excellence Best Practices (replaced 5 tips with 7 operational domains: CI/CD, IaC, Graviton, tagging, runbooks, multi-account governance); Cluster Segmentation (added deployment options decision framework — Serverless and EKS as segmentation alternatives); Incremental Data Processing (added migration path decision table, expanded to Delta Lake and Iceberg alongside Hudi); Data Catalog Migration (added catalog strategy decision table, modern Glue capabilities, trimmed duplicative config to cross-references); Ad Hoc Query Capabilities (added Iceberg support, Lake Formation, EMR Serverless for Trino, federated query). Cross-cutting modernization: Added references to EMR Serverless, EMR on EKS, S3 Tables, Apache Iceberg, Lake Formation, Trusted Identity Propagation, Graviton, Managed Scaling, Instance Fleets, MWAA, SageMaker Lakehouse, and CDK throughout. Removed deprecated content: EMRFS Consistent View, S3 Select, EMR Notebooks, Ganglia, Knox as primary gateway. Additional Resources refreshed (stale links removed, modern resources added). Contributors section restructured (2026 Update Contributors + Original Contributors). Appendix A expanded with Configuration and Dependencies section (12 new questions). Step-by-step configuration instructions trimmed and replaced with cross-references to official EMR documentation throughout. |