Publicado el — Deja un comentario

AWS Glue 6.0 delivers 30% price reduction and Iceberg v3 support

AWS Glue 6.0 is now generally available, delivering a 30% price reduction and introducing full support for Apache Iceberg v3, newer versions of Apache Hudi and Delta Lake, and new capabilities to improve developer productivity. AWS Glue 6.0 also upgrades runtime to Apache Spark 4.1, Python 3.13, and Scala 2.13.

With Apache Iceberg v3, AWS Glue 6.0 adds the VARIANT data type with automatic shredding for faster reads on semi-structured data, deletion vectors for high-performance row-level updates, geometry and geography data types for spatial processing, and flexible schema evolution through UNKNOWN data type and DEFAULT column values. Glue 6.0 also introduces features that boost developer productivity and performance, such as Spark Declarative Pipelines that eliminate repetitive orchestration code, Real-Time Mode streaming for sub-second latencies, and Arrow-native Python UDFs for improved PySpark performance. These capabilities help you implement large-scale ETL, recurring batch workloads, streaming analytics, and AI application development using AWS Glue.

AWS Glue 6.0 is available in all AWS Commercial, AWS GovCloud (US), and AWS China regions.

To get started, select Glue 6.0 from the version dropdown in the AWS Glue console or SageMaker Unified Studio when creating a new job, or migrate existing jobs using the Spark Upgrade Agent. To learn more, visit the AWS Glue documentation and AWS Glue pricing.

 

​AWS Glue 6.0 is now generally available, delivering a 30% price reduction and introducing full support for Apache Iceberg v3, newer versions of Apache Hudi and Delta Lake, and new capabilities to improve developer productivity. AWS Glue 6.0 also upgrades runtime to Apache Spark 4.1, Python 3.13, and Scala 2.13.
With Apache Iceberg v3, AWS Glue 6.0 adds the VARIANT data type with automatic shredding for faster reads on semi-structured data, deletion vectors for high-performance row-level updates, geometry and geography data types for spatial processing, and flexible schema evolution through UNKNOWN data type and DEFAULT column values. Glue 6.0 also introduces features that boost developer productivity and performance, such as Spark Declarative Pipelines that eliminate repetitive orchestration code, Real-Time Mode streaming for sub-second latencies, and Arrow-native Python UDFs for improved PySpark performance. These capabilities help you implement large-scale ETL, recurring batch workloads, streaming analytics, and AI application development using AWS Glue.
AWS Glue 6.0 is available in all AWS Commercial, AWS GovCloud (US), and AWS China regions.
To get started, select Glue 6.0 from the version dropdown in the AWS Glue console or SageMaker Unified Studio when creating a new job, or migrate existing jobs using the Spark Upgrade Agent. To learn more, visit the AWS Glue documentation and AWS Glue pricing.  

Publicado el — Deja un comentario

Amazon SES now supports open and click tracking override parameters

Amazon Simple Email Service (SES) now supports open and click tracking override parameters in the SendEmail and SendBulkEmail APIs. Senders can enable or disable open tracking and click tracking on an individual API call, rather than managing tracking preferences through separate configuration sets.

Previously, controlling tracking behavior required maintaining a distinct configuration set for each combination of open- and click-tracking settings. With this new capability, you specify the tracking preference directly in the send request, reducing configuration overhead and simplifying how you honor recipient-level tracking consent. This is useful for senders that must respect per-recipient consent choices to meet data protection requirements such as GDPR and CNIL guidance.

The tracking overrides apply per request and take precedence over the tracking behavior defined in the associated configuration set, giving you fine-grained control without changing your existing configuration set structure. There is no additional cost to use this feature.

This capability is available in all AWS Regions where Amazon SES is available.

To learn more, see the documentation on open and click tracking in the Amazon SES Developer Guide. 

 

​Amazon Simple Email Service (SES) now supports open and click tracking override parameters in the SendEmail and SendBulkEmail APIs. Senders can enable or disable open tracking and click tracking on an individual API call, rather than managing tracking preferences through separate configuration sets. Previously, controlling tracking behavior required maintaining a distinct configuration set for each combination of open- and click-tracking settings. With this new capability, you specify the tracking preference directly in the send request, reducing configuration overhead and simplifying how you honor recipient-level tracking consent. This is useful for senders that must respect per-recipient consent choices to meet data protection requirements such as GDPR and CNIL guidance. The tracking overrides apply per request and take precedence over the tracking behavior defined in the associated configuration set, giving you fine-grained control without changing your existing configuration set structure. There is no additional cost to use this feature. This capability is available in all AWS Regions where Amazon SES is available.
To learn more, see the documentation on open and click tracking in the Amazon SES Developer Guide.   

Publicado el — Deja un comentario

Amazon EC2 C8gd, M8gd and R8gd instances are now available in additional AWS Regions

Amazon Elastic Compute Cloud (Amazon EC2) C8gd, M8gd, and R8gd instances with up to 11.4 TB of local NVMe-based SSD block-level storage are now available in additional regions. C8gd instances are now available in Asia Pacific (Singapore), M8gd instances are available in Mexico (Central) and Asia Pacific (Melbourne), and R8gd instances are available in Europe (Zurich). These instances are powered by AWS Graviton4 processors, delivering up to 30% better performance over Graviton3-based instances. They have up to 40% higher performance for I/O intensive database workloads, and up to 20% faster query results for I/O intensive real-time data analytics than comparable AWS Graviton3-based instances. These instances are built on the AWS Nitro System and are a great fit for applications that need access to high-speed, low latency local storage.

Each instance is available in 12 different sizes. They provide up to 50 Gbps of network bandwidth and up to 40 Gbps of bandwidth to the Amazon Elastic Block Store (Amazon EBS). Additionally, customers can now adjust the network and Amazon EBS bandwidth on these instances by 25% using EC2 instance bandwidth weighting configuration, providing greater flexibility with the allocation of bandwidth resources to better optimize workloads. These instances offer Elastic Fabric Adapter (EFA) networking on 24xlarge, 48xlarge, metal-24xl, and metal-48xl sizes.

C8gd instances are ideal for compute-intensive workloads such as high-performance web servers, batch processing, distributed analytics, ad serving, video encoding, and gaming servers. M8gd instances are well-suited for balanced workloads including application servers, microservices, enterprise applications, and small to medium databases. R8gd instances are ideal for memory-intensive workloads such as in-memory databases, real-time big data analytics, large in-memory caches, and scientific computing applications.

To learn more, see Amazon C8gd Instances, Amazon M8gd Instances and Amazon R8gd Instances. To explore how to migrate your workloads to Graviton-based instances, see AWS Graviton Fast Start program and Porting Advisor for Graviton. To get started, see the AWS Management Console.

 

​Amazon Elastic Compute Cloud (Amazon EC2) C8gd, M8gd, and R8gd instances with up to 11.4 TB of local NVMe-based SSD block-level storage are now available in additional regions. C8gd instances are now available in Asia Pacific (Singapore), M8gd instances are available in Mexico (Central) and Asia Pacific (Melbourne), and R8gd instances are available in Europe (Zurich). These instances are powered by AWS Graviton4 processors, delivering up to 30% better performance over Graviton3-based instances. They have up to 40% higher performance for I/O intensive database workloads, and up to 20% faster query results for I/O intensive real-time data analytics than comparable AWS Graviton3-based instances. These instances are built on the AWS Nitro System and are a great fit for applications that need access to high-speed, low latency local storage. Each instance is available in 12 different sizes. They provide up to 50 Gbps of network bandwidth and up to 40 Gbps of bandwidth to the Amazon Elastic Block Store (Amazon EBS). Additionally, customers can now adjust the network and Amazon EBS bandwidth on these instances by 25% using EC2 instance bandwidth weighting configuration, providing greater flexibility with the allocation of bandwidth resources to better optimize workloads. These instances offer Elastic Fabric Adapter (EFA) networking on 24xlarge, 48xlarge, metal-24xl, and metal-48xl sizes.
C8gd instances are ideal for compute-intensive workloads such as high-performance web servers, batch processing, distributed analytics, ad serving, video encoding, and gaming servers. M8gd instances are well-suited for balanced workloads including application servers, microservices, enterprise applications, and small to medium databases. R8gd instances are ideal for memory-intensive workloads such as in-memory databases, real-time big data analytics, large in-memory caches, and scientific computing applications.
To learn more, see Amazon C8gd Instances, Amazon M8gd Instances and Amazon R8gd Instances. To explore how to migrate your workloads to Graviton-based instances, see AWS Graviton Fast Start program and Porting Advisor for Graviton. To get started, see the AWS Management Console.  

Publicado el — Deja un comentario

AWS announces the general availability of a new AWS Local Zone in Las Vegas, Nevada

AWS Local Zone in Las Vegas, Nevada is now generally available. The new AWS Local Zone supports Amazon Elastic Compute Cloud (Amazon EC2) C7i, M7i, R7i, and C8gn instances, Amazon Elastic Block Store (Amazon EBS) volume types gp3, gp2, io1, sc1, and st1, Amazon Elastic Container Service (Amazon ECS), Amazon Elastic Kubernetes Service (Amazon EKS), Application Load Balancer, and AWS Direct Connect.

AWS Local Zones are AWS infrastructure deployments that extend core services, such as compute, storage, networking, and other select services, closer to metropolitan areas worldwide. AWS Local Zones help you achieve single-digit millisecond latency for end-user workloads, meet data residency requirements, support AI/ML inference workloads, and accelerate migration and modernization of legacy applications to the cloud, all while maintaining consistent AWS APIs, tools, and services as AWS Regions. AWS Local Zones are available in more than 30 metropolitan areas worldwide.

To get started, enable the Las Vegas Local Zone (us-west-2-las-2a) from the Regions and Zones tab in the AWS Global View or by using the ModifyAvailabilityZoneGroup API. For pricing information, visit the AWS Local Zones pricing page. To learn more, visit the AWS Local Zones overview page.

 

​AWS Local Zone in Las Vegas, Nevada is now generally available. The new AWS Local Zone supports Amazon Elastic Compute Cloud (Amazon EC2) C7i, M7i, R7i, and C8gn instances, Amazon Elastic Block Store (Amazon EBS) volume types gp3, gp2, io1, sc1, and st1, Amazon Elastic Container Service (Amazon ECS), Amazon Elastic Kubernetes Service (Amazon EKS), Application Load Balancer, and AWS Direct Connect.
AWS Local Zones are AWS infrastructure deployments that extend core services, such as compute, storage, networking, and other select services, closer to metropolitan areas worldwide. AWS Local Zones help you achieve single-digit millisecond latency for end-user workloads, meet data residency requirements, support AI/ML inference workloads, and accelerate migration and modernization of legacy applications to the cloud, all while maintaining consistent AWS APIs, tools, and services as AWS Regions. AWS Local Zones are available in more than 30 metropolitan areas worldwide.
To get started, enable the Las Vegas Local Zone (us-west-2-las-2a) from the Regions and Zones tab in the AWS Global View or by using the ModifyAvailabilityZoneGroup API. For pricing information, visit the AWS Local Zones pricing page. To learn more, visit the AWS Local Zones overview page.  

Publicado el — Deja un comentario

Amazon Timestream for InfluxDB now supports customer managed keys

Amazon Timestream for InfluxDB now supports AWS Key Management Service (AWS KMS) customer managed keys for encrypting data at rest in InfluxDB 2 database instances, InfluxDB 2 Read Replicas, and InfluxDB 3 clusters. Customers select a symmetric AWS KMS key when creating a database resource.

Timestream for InfluxDB uses the selected key to encrypt the underlying database storage for InfluxDB 2 and InfluxDB 3 resources. The key must be in the same AWS account and AWS Region as the database resource. Customers specify the key during resource creation. The key cannot be changed after the resource is created.

Customer managed key support is available through the AWS Management Console, AWS Command Line Interface (AWS CLI), and Timestream for InfluxDB application programming interface (API). The feature is available in all AWS Regions where Timestream for InfluxDB is available. There is no additional Timestream for InfluxDB charge for using customer managed keys. Standard AWS KMS charges apply.

Support for Customer managed keys is available in all AWS Regions where Amazon Timestream for InfluxDB is available. To get started, open the Amazon Timestream console. For more information, see the Amazon Timestream for InfluxDB documentation and pricing page.

 

​Amazon Timestream for InfluxDB now supports AWS Key Management Service (AWS KMS) customer managed keys for encrypting data at rest in InfluxDB 2 database instances, InfluxDB 2 Read Replicas, and InfluxDB 3 clusters. Customers select a symmetric AWS KMS key when creating a database resource.
Timestream for InfluxDB uses the selected key to encrypt the underlying database storage for InfluxDB 2 and InfluxDB 3 resources. The key must be in the same AWS account and AWS Region as the database resource. Customers specify the key during resource creation. The key cannot be changed after the resource is created.
Customer managed key support is available through the AWS Management Console, AWS Command Line Interface (AWS CLI), and Timestream for InfluxDB application programming interface (API). The feature is available in all AWS Regions where Timestream for InfluxDB is available. There is no additional Timestream for InfluxDB charge for using customer managed keys. Standard AWS KMS charges apply. Support for Customer managed keys is available in all AWS Regions where Amazon Timestream for InfluxDB is available. To get started, open the Amazon Timestream console. For more information, see the Amazon Timestream for InfluxDB documentation and pricing page.  

Publicado el — Deja un comentario

Amazon EC2 P6-B300 instances are now available in the Asia Pacific (Seoul) Region

Starting today, Amazon Elastic Cloud Compute (Amazon EC2) P6-B300 instances are available in the Asia Pacific (Seoul) Region. P6-B300 instances provide 8xNVIDIA Blackwell Ultra GPUs with 2.1 TB high bandwidth GPU memory, 6.4 Tbps EFA networking, 300 Gbps dedicated ENA throughput, and 4 TB of system memory.

P6-B300 instances deliver 2x networking bandwidth, 1.5x GPU memory size, and 1.5x GPU TFLOPS (at FP4, without sparsity) compared to P6-B200 instances, making them well suited to train and deploy large trillion-parameter foundation models (FMs) and large language models (LLMs) with sophisticated techniques. The higher networking and larger memory deliver faster training times and more token throughput for AI workloads.

P6-B300 instances are now available in p6-b300.48xlarge size in the following AWS Regions: US West (Oregon), AWS GovCloud (US-East), US East (N. Virginia) and Asia Pacific (Seoul). To learn more about P6-B300 instances, visit Amazon EC2 P6 instances.

 

​Starting today, Amazon Elastic Cloud Compute (Amazon EC2) P6-B300 instances are available in the Asia Pacific (Seoul) Region. P6-B300 instances provide 8xNVIDIA Blackwell Ultra GPUs with 2.1 TB high bandwidth GPU memory, 6.4 Tbps EFA networking, 300 Gbps dedicated ENA throughput, and 4 TB of system memory. P6-B300 instances deliver 2x networking bandwidth, 1.5x GPU memory size, and 1.5x GPU TFLOPS (at FP4, without sparsity) compared to P6-B200 instances, making them well suited to train and deploy large trillion-parameter foundation models (FMs) and large language models (LLMs) with sophisticated techniques. The higher networking and larger memory deliver faster training times and more token throughput for AI workloads. P6-B300 instances are now available in p6-b300.48xlarge size in the following AWS Regions: US West (Oregon), AWS GovCloud (US-East), US East (N. Virginia) and Asia Pacific (Seoul). To learn more about P6-B300 instances, visit Amazon EC2 P6 instances.  

Publicado el — Deja un comentario

Amazon EKS now supports certificate authority (CA) rotation with automated lifecycle management

Today, Amazon Elastic Kubernetes Service (Amazon EKS) announced certificate authority (CA) rotation, enabling customers to rotate their cluster’s CA through a managed lifecycle with automated safeguards. Each Amazon EKS cluster has its own CA that allows encrypted connections to the cluster’s Kubernetes API, and now you can rotate the CA before it expires to ensure your cluster remains operational and secure.

Amazon EKS clusters created since launch in 2018 have CAs with a 10-year validity period, and clusters from that era are now approaching the point where CA rotation activities should begin. CA rotation in Amazon EKS is a shared responsibility. Amazon EKS manages the rotation lifecycle and automatically updates AWS-managed components to trust the successor CA. Customers are responsible for replacing their worker nodes and updating external clients to trust the successor CA before it is activated. EKS Auto Mode instances and AWS Fargate nodes are updated automatically by AWS, but customers are still responsible for updating any external clients that connect to the cluster’s API server. Amazon EKS provides automated safeguards to support customers through this process, including advance notifications before CA expiration, automatic appending of a successor CA if one is not created by the customer, and automatic activation if the customer does not activate on their own schedule. A rollback capability allows customers to revert to the previous CA to resolve any issues that may arise with their updates during the transition to the successor CA.

Amazon EKS CA rotation is available at no additional cost in all commercial AWS Regions. To get started with CA rotation, you can use the AWS CLI, EKS APIs, CloudFormation, and the AWS console. For more information, see the Amazon EKS documentation and Deep dive into Amazon EKS certificate authority rotation.

 

​Today, Amazon Elastic Kubernetes Service (Amazon EKS) announced certificate authority (CA) rotation, enabling customers to rotate their cluster’s CA through a managed lifecycle with automated safeguards. Each Amazon EKS cluster has its own CA that allows encrypted connections to the cluster’s Kubernetes API, and now you can rotate the CA before it expires to ensure your cluster remains operational and secure.
Amazon EKS clusters created since launch in 2018 have CAs with a 10-year validity period, and clusters from that era are now approaching the point where CA rotation activities should begin. CA rotation in Amazon EKS is a shared responsibility. Amazon EKS manages the rotation lifecycle and automatically updates AWS-managed components to trust the successor CA. Customers are responsible for replacing their worker nodes and updating external clients to trust the successor CA before it is activated. EKS Auto Mode instances and AWS Fargate nodes are updated automatically by AWS, but customers are still responsible for updating any external clients that connect to the cluster’s API server. Amazon EKS provides automated safeguards to support customers through this process, including advance notifications before CA expiration, automatic appending of a successor CA if one is not created by the customer, and automatic activation if the customer does not activate on their own schedule. A rollback capability allows customers to revert to the previous CA to resolve any issues that may arise with their updates during the transition to the successor CA.
Amazon EKS CA rotation is available at no additional cost in all commercial AWS Regions. To get started with CA rotation, you can use the AWS CLI, EKS APIs, CloudFormation, and the AWS console. For more information, see the Amazon EKS documentation and Deep dive into Amazon EKS certificate authority rotation.  

Publicado el — Deja un comentario

Amazon CloudFront now supports Origin Access Control (OAC) for Amazon S3 Multi-Region Access Points

Starting today, customers can protect their origins using Amazon S3 Multi-Region Access Points (MRAP) by using CloudFront Origin Access Control (OAC) to only allow access from designated CloudFront distributions.

Customers use Amazon S3 MRAP with CloudFront to serve content from a single global endpoint that automatically routes to the closest available replicated bucket across regions during a cache miss, improving performance and resilience for globally distributed users. Previously, customers had to compute and forward their own Asymmetric Signature Version 4 (SigV4a) Authorization header using a custom Lambda@Edge Function. Now, CloudFront natively signs requests to S3 MRAP origins. Customers get faster cache-miss fills from the nearest region and restricted, OAC-secured MRAP access without  custom Authorization header computation.

CloudFront OAC support for Amazon S3 MRAP origins is available worldwide, except in the CloudFront China region. To get started, use the CloudFront Console, SDK, CLI, or CloudFormation to enable OAC when configuring your Amazon S3 MRAP endpoint with CloudFront. For more information, refer to the CloudFront Developer Guide. There are no additional fees associated with this feature

 

​Starting today, customers can protect their origins using Amazon S3 Multi-Region Access Points (MRAP) by using CloudFront Origin Access Control (OAC) to only allow access from designated CloudFront distributions.
Customers use Amazon S3 MRAP with CloudFront to serve content from a single global endpoint that automatically routes to the closest available replicated bucket across regions during a cache miss, improving performance and resilience for globally distributed users. Previously, customers had to compute and forward their own Asymmetric Signature Version 4 (SigV4a) Authorization header using a custom Lambda@Edge Function. Now, CloudFront natively signs requests to S3 MRAP origins. Customers get faster cache-miss fills from the nearest region and restricted, OAC-secured MRAP access without  custom Authorization header computation.
CloudFront OAC support for Amazon S3 MRAP origins is available worldwide, except in the CloudFront China region. To get started, use the CloudFront Console, SDK, CLI, or CloudFormation to enable OAC when configuring your Amazon S3 MRAP endpoint with CloudFront. For more information, refer to the CloudFront Developer Guide. There are no additional fees associated with this feature  

Publicado el — Deja un comentario

AWS Partner Central agents MCP Server now supports OAuth with AWS Sign-In

AWS partners can now access AWS Partner Central agents from tools they already use, such as Amazon Quick and Kiro, using OAuth through AWS Sign-In. Partners can authorize agent access with their existing AWS identities, sign-in methods, IAM permissions, and governance controls without installing or maintaining additional authentication software.

Previously, AWS Partners needed to set up an MCP proxy with SigV4 credentials to access Partner Central agents from their existing tools, or sign in to AWS Partner Central through the AWS Management Console with IAM credentials. OAuth simplifies this by allowing partners to use AWS Sign-In to authorize tools such as Amazon Quick and Kiro to access Partner Central agents. Partners can use OAuth from their existing tools for co-sell engagements, AWS funding applications, and AWS Marketplace seller setup. Administrators can govern access with IAM policies, global condition keys, token introspection and revocation APIs, dynamic client registration, and CloudTrail audit events.

OAuth support is available to AWS Partners through AWS Partner Central agents MCP Server, which is available in the US East (N. Virginia) Region. To learn more, visit  Getting started with the Partner Central agents MCP Server , and  Sign-In with OAuth 2.0 .

 

​AWS partners can now access AWS Partner Central agents from tools they already use, such as Amazon Quick and Kiro, using OAuth through AWS Sign-In. Partners can authorize agent access with their existing AWS identities, sign-in methods, IAM permissions, and governance controls without installing or maintaining additional authentication software.
Previously, AWS Partners needed to set up an MCP proxy with SigV4 credentials to access Partner Central agents from their existing tools, or sign in to AWS Partner Central through the AWS Management Console with IAM credentials. OAuth simplifies this by allowing partners to use AWS Sign-In to authorize tools such as Amazon Quick and Kiro to access Partner Central agents. Partners can use OAuth from their existing tools for co-sell engagements, AWS funding applications, and AWS Marketplace seller setup. Administrators can govern access with IAM policies, global condition keys, token introspection and revocation APIs, dynamic client registration, and CloudTrail audit events.
OAuth support is available to AWS Partners through AWS Partner Central agents MCP Server, which is available in the US East (N. Virginia) Region. To learn more, visit  Getting started with the Partner Central agents MCP Server , and  Sign-In with OAuth 2.0 .  

Publicado el — Deja un comentario

Generative AI Inference Recommendation for Amazon SageMaker now available in the SageMaker AI Studio

Amazon SageMaker AI now offers Generative AI Inference Recommendations in SageMaker AI Studio, giving customers a guided, low-code, no-code path to find the best inference configuration for their workload. This builds on the API-based launch in April 2026, extending the same benchmarking infrastructure to teams that prefer a visual workflow over programmatic access.

Deploying generative AI models in production requires finding the right combination of instance type, serving container, and optimization strategy. Getting this right typically involves weeks of manual benchmarking, configuration tuning, and trial-and-error, with no easy way to know if the final setup is actually optimal. With the new experience, customers describe their workload and what matters most, whether that’s latency, throughput, or cost, and SageMaker AI does the rest. It benchmarks multiple configurations on real GPU infrastructure using NVIDIA AIPerf, applies goal-aligned techniques like speculative decoding for throughput or kernel tuning for latency, and returns ranked, production-ready recommendations with measured performance data. Teams get to a validated configuration in hours instead of weeks, without needing to decide which techniques to apply or how to configure them.

With the new experience, customers describe their workload and what matters most, whether that’s latency, throughput, or cost, and SageMaker AI does the rest. It benchmarks multiple configurations on real GPU infrastructure using NVIDIA AIPerf, applies goal-aligned techniques like speculative decoding for throughput or kernel tuning for latency, and returns ranked, production-ready recommendations with measured performance data. Teams get to a validated configuration in hours instead of weeks, without needing to decide which techniques to apply or how to configure them.

In SageMaker AI Studio under Jobs, Inference optimization, customers select a use-case profile (Interact, Generate, Summarize, or Custom), choose an optimization goal (minimize latency, maximize throughput, or minimize cost), and pick their model from JumpStart, S3, Model Registry, or an existing SageMaker model. Recommendations are ranked by TTFT, inter-token latency, throughput, and cost, and can be compared visually before deploying to a SageMaker real-time endpoint directly from Studio.

There is no additional cost for generating recommendations. Standard compute costs apply for optimization jobs and endpoints provisioned during benchmarking. This capability is available in US East (N. Virginia), US West (Oregon), US East (Ohio), Europe (Ireland), Europe (Frankfurt), Asia Pacific (Singapore), Asia Pacific (Tokyo). To learn more, visit the blog post or the documentation.

 

​Amazon SageMaker AI now offers Generative AI Inference Recommendations in SageMaker AI Studio, giving customers a guided, low-code, no-code path to find the best inference configuration for their workload. This builds on the API-based launch in April 2026, extending the same benchmarking infrastructure to teams that prefer a visual workflow over programmatic access.
Deploying generative AI models in production requires finding the right combination of instance type, serving container, and optimization strategy. Getting this right typically involves weeks of manual benchmarking, configuration tuning, and trial-and-error, with no easy way to know if the final setup is actually optimal. With the new experience, customers describe their workload and what matters most, whether that’s latency, throughput, or cost, and SageMaker AI does the rest. It benchmarks multiple configurations on real GPU infrastructure using NVIDIA AIPerf, applies goal-aligned techniques like speculative decoding for throughput or kernel tuning for latency, and returns ranked, production-ready recommendations with measured performance data. Teams get to a validated configuration in hours instead of weeks, without needing to decide which techniques to apply or how to configure them.
With the new experience, customers describe their workload and what matters most, whether that’s latency, throughput, or cost, and SageMaker AI does the rest. It benchmarks multiple configurations on real GPU infrastructure using NVIDIA AIPerf, applies goal-aligned techniques like speculative decoding for throughput or kernel tuning for latency, and returns ranked, production-ready recommendations with measured performance data. Teams get to a validated configuration in hours instead of weeks, without needing to decide which techniques to apply or how to configure them.
In SageMaker AI Studio under Jobs, Inference optimization, customers select a use-case profile (Interact, Generate, Summarize, or Custom), choose an optimization goal (minimize latency, maximize throughput, or minimize cost), and pick their model from JumpStart, S3, Model Registry, or an existing SageMaker model. Recommendations are ranked by TTFT, inter-token latency, throughput, and cost, and can be compared visually before deploying to a SageMaker real-time endpoint directly from Studio.
There is no additional cost for generating recommendations. Standard compute costs apply for optimization jobs and endpoints provisioned during benchmarking. This capability is available in US East (N. Virginia), US West (Oregon), US East (Ohio), Europe (Ireland), Europe (Frankfurt), Asia Pacific (Singapore), Asia Pacific (Tokyo). To learn more, visit the blog post or the documentation.