Publicado el — Deja un comentario

Amazon Quick agentic AI capabilities are now available in AWS GovCloud (US-West)

Today, AWS announces that Amazon Quick’s agentic AI capabilities are now available in AWS GovCloud (US-West), bringing an agentic AI teammate to government and regulated-industry teams within an isolated, FedRAMP Class D (formerly High) authorized environment. Building on the analytics and business intelligence capabilities already available to customers in AWS GovCloud (US), Quick now turns questions into actions, helping teams drive mission-critical decisions faster without switching applications.

With this launch, teams can build custom chat agents tailored to mission-specific workflows — including procurement, ATO compliance, and grants management — while keeping data hosted and processed entirely within the AWS GovCloud (US-West) Region. Spaces enforce least-privilege access by scoping information to the appropriate program office or mission area, ensuring analysts only access mission-relevant data. Quick also integrates with tools teams already rely on, including Microsoft 365, SharePoint, and OneDrive via GCC High connectors, as well as browser extensions.

AWS GovCloud (US) Regions are isolated AWS Regions operated by U.S. citizens on U.S. soil, purpose-built to host sensitive data and regulated workloads. Customers can address the most stringent U.S. government security and compliance requirements, including the FedRAMP Class D (formerly High) baseline, Department of Defense Cloud Computing Security Requirements Guide (DoD SRG) Impact Levels 4 and 5, International Traffic in Arms Regulations (ITAR), Criminal Justice Information Services (CJIS), and Federal Information Processing Standard (FIPS) 140-3. Inference on authorized foundation models is processed within the AWS GovCloud (US-West) Region, and enterprise governance features are available at launch.

With this launch, Amazon Quick’s agentic AI capabilities are available in 8 AWS Regions: US East (N. Virginia), US West (Oregon), Europe (Frankfurt, Ireland, London), Asia Pacific (Sydney, Tokyo), and AWS GovCloud (US-West). 

To learn more, visit the Amazon Quick product page and AWS GovCloud (US) documentation

 

​Today, AWS announces that Amazon Quick’s agentic AI capabilities are now available in AWS GovCloud (US-West), bringing an agentic AI teammate to government and regulated-industry teams within an isolated, FedRAMP Class D (formerly High) authorized environment. Building on the analytics and business intelligence capabilities already available to customers in AWS GovCloud (US), Quick now turns questions into actions, helping teams drive mission-critical decisions faster without switching applications.
With this launch, teams can build custom chat agents tailored to mission-specific workflows — including procurement, ATO compliance, and grants management — while keeping data hosted and processed entirely within the AWS GovCloud (US-West) Region. Spaces enforce least-privilege access by scoping information to the appropriate program office or mission area, ensuring analysts only access mission-relevant data. Quick also integrates with tools teams already rely on, including Microsoft 365, SharePoint, and OneDrive via GCC High connectors, as well as browser extensions.
AWS GovCloud (US) Regions are isolated AWS Regions operated by U.S. citizens on U.S. soil, purpose-built to host sensitive data and regulated workloads. Customers can address the most stringent U.S. government security and compliance requirements, including the FedRAMP Class D (formerly High) baseline, Department of Defense Cloud Computing Security Requirements Guide (DoD SRG) Impact Levels 4 and 5, International Traffic in Arms Regulations (ITAR), Criminal Justice Information Services (CJIS), and Federal Information Processing Standard (FIPS) 140-3. Inference on authorized foundation models is processed within the AWS GovCloud (US-West) Region, and enterprise governance features are available at launch.
With this launch, Amazon Quick’s agentic AI capabilities are available in 8 AWS Regions: US East (N. Virginia), US West (Oregon), Europe (Frankfurt, Ireland, London), Asia Pacific (Sydney, Tokyo), and AWS GovCloud (US-West). 
To learn more, visit the Amazon Quick product page and AWS GovCloud (US) documentation  

Publicado el — Deja un comentario

Amazon EC2 R8a instances are now available in Canada (Central) region

Starting today, Amazon EC2 R8a instances are now available in Canada (Central) Region. These instances, feature 5th Gen AMD EPYC processors (formerly code named Turin) with a maximum frequency of 4.5 GHz, deliver up to 30% higher performance, and up to 19% better price-performance compared to R7a instances.

R8a instances deliver 45% more memory bandwidth compared to R7a instances, making these instances ideal for latency sensitive workloads. Compared to Amazon EC2 R7a instances, R8a instances provide up to 60% faster performance for GroovyJVM, allowing higher request throughput and better response times for business-critical applications.

Built on the AWS Nitro System using sixth generation Nitro Cards, R8a instances are ideal for high performance, memory-intensive workloads, such as SQL and NoSQL databases, distributed web scale in-memory caches, in-memory databases, real-time big data analytics, and Electronic Design Automation (EDA) applications. R8a instances offer 12 sizes including 2 bare metal sizes. Amazon EC2 R8a instances are SAP-certified, and providing 38% more SAPS compared to R7a instances.

To get started, sign in to the AWS Management Console. For more information about the new instances, visit the Amazon EC2 R8a instance page. 

 

​Starting today, Amazon EC2 R8a instances are now available in Canada (Central) Region. These instances, feature 5th Gen AMD EPYC processors (formerly code named Turin) with a maximum frequency of 4.5 GHz, deliver up to 30% higher performance, and up to 19% better price-performance compared to R7a instances.
R8a instances deliver 45% more memory bandwidth compared to R7a instances, making these instances ideal for latency sensitive workloads. Compared to Amazon EC2 R7a instances, R8a instances provide up to 60% faster performance for GroovyJVM, allowing higher request throughput and better response times for business-critical applications.
Built on the AWS Nitro System using sixth generation Nitro Cards, R8a instances are ideal for high performance, memory-intensive workloads, such as SQL and NoSQL databases, distributed web scale in-memory caches, in-memory databases, real-time big data analytics, and Electronic Design Automation (EDA) applications. R8a instances offer 12 sizes including 2 bare metal sizes. Amazon EC2 R8a instances are SAP-certified, and providing 38% more SAPS compared to R7a instances.
To get started, sign in to the AWS Management Console. For more information about the new instances, visit the Amazon EC2 R8a instance page.   

Publicado el — Deja un comentario

Amazon Bedrock expands IAM principal cost allocation to the bedrock-mantle endpoint

Amazon Bedrock is a fully managed service that provides secure, enterprise-grade access to high-performing foundation models from leading AI companies, enabling you to build and scale generative AI applications. Amazon Bedrock now supports cost allocation by AWS Identity and Access Management (IAM) principal, including IAM users and roles, for model inference requests made through the bedrock-mantle endpoint. This extends the capability previously available for the bedrock-runtime endpoint, helping customers attribute inference costs across users, teams, projects, and applications.

Customers can tag IAM users and roles with attributes such as team, project, or cost center, activate them as cost allocation tags, and analyze bedrock-mantle inference costs by those tags in AWS Cost Explorer or at the line-item level in AWS Cost and Usage Report 2.0 (CUR 2.0). To get started, activate your IAM principal tags in the AWS Billing and Cost Management console. Then filter or group costs by those tags in Cost Explorer, or create a CUR 2.0 data export and select Include caller identity (IAM principal) allocation data.

This feature is available in all AWS Regions where the bedrock-mantle endpoint is available. To learn more, see Using IAM principal for cost allocation and IAM principal attribution in Amazon Bedrock.

 

​Amazon Bedrock is a fully managed service that provides secure, enterprise-grade access to high-performing foundation models from leading AI companies, enabling you to build and scale generative AI applications. Amazon Bedrock now supports cost allocation by AWS Identity and Access Management (IAM) principal, including IAM users and roles, for model inference requests made through the bedrock-mantle endpoint. This extends the capability previously available for the bedrock-runtime endpoint, helping customers attribute inference costs across users, teams, projects, and applications.
Customers can tag IAM users and roles with attributes such as team, project, or cost center, activate them as cost allocation tags, and analyze bedrock-mantle inference costs by those tags in AWS Cost Explorer or at the line-item level in AWS Cost and Usage Report 2.0 (CUR 2.0). To get started, activate your IAM principal tags in the AWS Billing and Cost Management console. Then filter or group costs by those tags in Cost Explorer, or create a CUR 2.0 data export and select Include caller identity (IAM principal) allocation data.
This feature is available in all AWS Regions where the bedrock-mantle endpoint is available. To learn more, see Using IAM principal for cost allocation and IAM principal attribution in Amazon Bedrock.  

Publicado el — Deja un comentario

AWS Secrets Manager adds managed external secrets support for Jenkins and SonarQube

AWS Secrets Manager now extends its managed external secrets capability to include Jenkins API Tokens and SonarQube Tokens, enabling you to automatically rotate these third-party credentials directly from the AWS console without writing any custom rotation code.

For Jenkins, Secrets Manager mints a new token and revokes the old one only after the replacement is verified active, so your continuous integration and continuous delivery (CI/CD) jobs transition without interruption. Rotation supports both self-rotation, where the token being rotated authenticates its own replacement, and admin-assisted rotation, where a separate admin token performs the generate and revoke operations. For SonarQube, you can rotate three types of tokens — User Tokens, Global Analysis Tokens, and Project Analysis Tokens — via SonarQube’s Web API. User Tokens support self-rotation, while analysis tokens are rotated using an admin token.

These integrations join existing managed external secrets support for BigID, Confluent Cloud, Datadog, GitLab, MongoDB Atlas, Okta, Paddle, Salesforce, and Snowflake.

Jenkins and SonarQube managed external secrets are available in all AWS Regions where AWS Secrets Manager managed external secrets is supported. To learn more, visit the  AWS Secrets Manager managed external secrets documentation .

 

​AWS Secrets Manager now extends its managed external secrets capability to include Jenkins API Tokens and SonarQube Tokens, enabling you to automatically rotate these third-party credentials directly from the AWS console without writing any custom rotation code.
For Jenkins, Secrets Manager mints a new token and revokes the old one only after the replacement is verified active, so your continuous integration and continuous delivery (CI/CD) jobs transition without interruption. Rotation supports both self-rotation, where the token being rotated authenticates its own replacement, and admin-assisted rotation, where a separate admin token performs the generate and revoke operations. For SonarQube, you can rotate three types of tokens — User Tokens, Global Analysis Tokens, and Project Analysis Tokens — via SonarQube’s Web API. User Tokens support self-rotation, while analysis tokens are rotated using an admin token.
These integrations join existing managed external secrets support for BigID, Confluent Cloud, Datadog, GitLab, MongoDB Atlas, Okta, Paddle, Salesforce, and Snowflake.
Jenkins and SonarQube managed external secrets are available in all AWS Regions where AWS Secrets Manager managed external secrets is supported. To learn more, visit the  AWS Secrets Manager managed external secrets documentation .  

Publicado el — Deja un comentario

Optimización de la curva de rendimiento Frontier

Optimización de la curva de rendimiento Frontier

Nueve íconos de aplicaciones, incluidos Correo, JetBrains, GitHub, PowerPoint, Excel, Visual Studio Code y Microsoft OneDrive, dispuestos en fila sobre un fondo degradado verde y rosa desenfocado con líneas punteadas.

Por: Mustafa Suleyman.

El Tokenmaxxing (maximizar la eficiencia de cada token en el modelo) ha sido la historia de los últimos meses, pero la eficiencia de tokens es el siguiente gran foco en la industria. ¿Cómo conseguimos el mejor rendimiento posible por token invertido y el mejor resultado real para el cliente por cada dólar invertido?

Para construir una Frontier Firm, hay que optimizar el rendimiento de la frontera frente al coste. Elegir dónde quieren situarse en esa curva es fundamental. A través de optimizar de manera conjunta sus modelos, harnesses (herramientas de prueba) y entornos de ejecución (RLEs) pueden elegir un punto de la curva que se adapte a tu empresa.

En la mayoría de los casos, los modelos Frontier generalistas no son necesarios para todas las tareas. Al ajustar modelos para un producto específico, pueden mantener o incluso superar el rendimiento de frontera, mientras reducen de manera importante los costes de tokens.

Aquí es donde hemos centrado nuestra máquina de subida de colinas MAI durante el último trimestre, y los resultados son bastante buenos. En estos días hemos lanzado MAI-Cyber-1-Flash optimizado para nuestro harness MDASH.

En conjunto, el sistema se situó en el número 1 del referente líder de CyberGym – para superar a Mythos por 12 puntos – a un 50% del coste. Y, de manera sorprendente, también lo servimos en los H100.

Fue diseñado para gestionar hasta el 90% de las tareas de manera eficiente, de modo que MDASH pueda reservar los modelos más grandes y caros de nuestra flota (en este caso GPT 5.4) para el 10% de problemas con una dificultad excepcional que en verdad los necesitan.

Como mencionó Satya en nuestra llamada de resultados del cuarto trimestre, desde el trimestre pasado hemos lanzado más de una docena de nuevos modelos en imagen, voz, transcripción, programación y seguridad, y ya impulsan muchos de los productos más utilizados de Microsoft para mantener o mejorar la calidad mientras usan de manera significativa menos tokens, en muchos casos para ahorrar entre un 50 y un 90% del coste de la GPU:

  • Construimos MAI-Code-1-Flash de la mano con nuestros colegas de GitHub, donde desde junio millones de desarrolladores lo han utilizado en su trabajo diario. Un 10% más tasa de aceptación de código y un 10% menor en el uso mediano de tokens que GPT-5.4 Mini y Claude Haiku 4.5 en VS Code, para ya mostrar una mejor retención.
  • Luego entrenamos ese mismo punto de control dentro de un entorno Excel RL para lograr un rendimiento comparable al de GPT-5.6 para las tareas más comunes, siendo más rentables y lo suficientemente pequeño para servir en un A100 o H100 frente a los aceleradores más modernos y caros.
  • MAI-Image-2.5-Flash es ahora el predeterminado de extremo a extremo en Bing Image Creator, en producción en PowerPoint, donde reduce los costes de la GPU hasta un 84% en comparación con GPT-Image-2, y es el predeterminado para escenarios clave de edición de OneDrive, donde ha aumentado las tasas de guardado en un 26% y ofrece hasta 2,5 veces mayor eficiencia de tokens.
  • MAI-Voice-2-Flash ahora impulsa el Centro de Contacto Dynamics 365, donde clientes como T-Mobile y EasyJet desarrollan sus agentes de call center, para reducir los costes de la GPU hasta en un 89%.
  • MAI-Transcribe-1.5 ahora cubre el flujo de trabajo multilingüe de Dragon Copilot en 58 idiomas — una solución utilizada por 170.000 proveedores médicos que procesaron 28 millones de encuentros con pacientes el trimestre pasado, donde nuestras pruebas muestran una reducción relativa del 50% en las tasas de error en transcripción e identificación del idioma.

Y además, al diseñar en conjunto nuestros modelos con nuestro propio silicio, vemos un 40% más de rendimiento por vatio al ejecutar modelos MAI en Maia 200.

Pero el beneficio no es solo el coste. Es la resiliencia. Ahora toda empresa debe asumir que cualquier modelo del que depende podría desaparecer, ya sea por un incidente de seguridad, una desalineación empresarial o de política, o un cambio geopolítico.

Cada modelo en un producto o sistema agéntico debería ser sustituible, y eso solo es posible cuando construyen el harness, el contexto, la memoria y el espacio de acciones de manera independiente de una sola familia de modelos. Esa es la máquina de subir cuestas que hemos construido.

Creemos que esto es el comienzo de una curva de rendimiento en verdad nueva. Su forma representa un sistema más que un modelo, y recorrer esta curva ofrece mejor calidad, menor coste y más variedad.

Este ha sido un verano de trabajo duro pero maravilloso por parte del equipo. Somos muy conscientes de lo temprano que es esto y de cuánto aún nos queda por aprender. Pero la dirección es clara, subimos la cuesta para avanzar en la frontera de la curva coste-resultado, y compartiremos lo que aprendamos en el camino. Queda mucho más por venir.

The post Optimización de la curva de rendimiento Frontier appeared first on Source LATAM.

 

​The post Optimización de la curva de rendimiento Frontier appeared first on Source LATAM.  

Publicado el — Deja un comentario

Ivonne Mejía, nueva directora general de Microsoft México


megaphone

Ivonne Mejía, nueva directora general de Microsoft México

Persona posa detrás de un logo de Microsoft con los brazos cruzados

Ciudad de México – Microsoft anuncia el nombramiento de Ivonne Mejía como nueva directora general de Microsoft México a partir del 1 de septiembre de 2026. En esta posición, liderará la estrategia para crear valor para los clientes, acelerar el crecimiento de Microsoft en México y fortalecer las alianzas con clientes y socios de negocio. Durante el mes de agosto, Ivonne colaborará estrechamente con Rafael Sánchez, quien estuvo al frente de Microsoft México desde 2022, para asegurar una transición ordenada y fluida en la dirección general de la compañía en el país.

Con más de 15 años de trayectoria en Microsoft, Ivonne aporta una sólida experiencia y un profundo conocimiento del negocio, construidos a través de posiciones de liderazgo en áreas estratégicas de la compañía. A lo largo de su carrera, ha encabezado iniciativas de transformación digital e innovación, acelerado resultados de negocio y consolidado relaciones de confianza con clientes y socios, combinando visión estratégica, capacidad de ejecución y un firme compromiso con el crecimiento del ecosistema tecnológico.

Persona en un pasillo con una camisa y sosteniendo un saco en un brazo

“Estoy profundamente agradecida por esta oportunidad y por la confianza. Creo en el potencial de nuestro país y en la gente que lo hace posible. Asumo este rol con una convicción clara: nuestro crecimiento vendrá de nuestra gente y de nuestros clientes y socios de negocio. Mi enfoque será construir juntos, pensar en grande y aprovechar al máximo la fuerza de todo Microsoft para servir mejor a México”.

El liderazgo de Ivonne se distingue por su capacidad para unir perspectivas, potenciar el talento y transformar equipos diversos en equipos de alto desempeño. Esta combinación de visión estratégica, cercanía y enfoque en las personas será fundamental para conducir a Microsoft México hacia su siguiente capítulo, acelerar el crecimiento de la compañía y ampliar el valor que genera para sus clientes, socios de negocio y el ecosistema tecnológico del país.

La trayectoria de Ivonne combina una sólida formación académica con una valiosa experiencia en los ámbitos tecnológico y comercial. Es ingeniera electrónica por la Universidad Nacional de Colombia, cuenta con un MBA del Tecnológico de Monterrey y cursó un Diplomado en Administración y Dirección de Empresas en el IPADE. Antes de incorporarse a Microsoft en 2010, desarrolló su carrera en Rockwell Automation, donde adquirió una visión integral del negocio y fortaleció su experiencia comercial, sentando las bases de un liderazgo que conecta tecnología, estrategia y cercanía con los clientes.

El legado de Rafael Sánchez: transformación, crecimiento e innovación en México

Después de más de cuatro años al frente de Microsoft México, Rafa Sánchez ha decidido cerrar este capítulo de su trayectoria para emprender un nuevo proyecto profesional fuera de Microsoft. Durante su gestión como director general, impulsó el crecimiento de la organización en el país y fortaleció su presencia en la industria tecnológica, dejando un legado significativo para la compañía, sus clientes, socios de negocio y el ecosistema digital en México. Así mismo, impulsó una cultura basada en la colaboración, la inclusión y la excelencia, y lideró iniciativas clave para la organización.

Bajo su liderazgo se consolidó la transformación de las operaciones en el país, se fortalecieron las capacidades comerciales, se inauguró la Región de Centros de Datos de Querétaro y se relanzó el Microsoft Innovation Hub. También promovió iniciativas de diversidad, desarrollo de talento y representación de la industria, contribuyendo a posicionar a México como un referente dentro de Microsoft y a preparar al país para aprovechar las oportunidades de la era de la IA.

Con más de 30 años de experiencia en la industria de tecnologías de la información, Rafael iniciará una nueva etapa en su trayectoria profesional dentro de la industria de TI, que se dará a conocer próximamente.

El nombramiento de Ivonne Mejía marca un nuevo capítulo para Microsoft México y un hito en la historia de la compañía en el país: la primera mujer en asumir la dirección general, desde donde liderará esta nueva etapa de crecimiento, innovación e impacto.

###

Acerca de Microsoft

Microsoft (Nasdaq «MSFT» @microsoft) crea plataformas y herramientas impulsadas por la IA para ofrecer soluciones innovadoras que satisfagan las necesidades cambiantes de nuestros clientes. La empresa de tecnología está comprometida con hacer que la IA esté ampliamente disponible, de manera responsable, con la misión de empoderar a cada persona y a cada organización en el planeta para lograr más.

Contacto de prensa:   

Microsoft                                            Assembly México    

Tere Rodríguez                     microsoftMexico@assemblyinc.com     

teresar@microsoft.com        55 5350 1500    

The post Ivonne Mejía, nueva directora general de Microsoft México appeared first on Source LATAM.

 

​The post Ivonne Mejía, nueva directora general de Microsoft México appeared first on Source LATAM.  

Publicado el — Deja un comentario

LocateAnything-3B, Qwen-AgentWorld-35B-A3B, and Qwen3.5-122B-A10B models now available on Amazon SageMaker JumpStart

NVIDIA’s LocateAnything-3B, Qwen’s Qwen-AgentWorld-35B-A3B, and Qwen’s Qwen3.5-122B-A10B models are now available on Amazon SageMaker JumpStart, expanding the portfolio of foundation models available to AWS customers. These three models bring specialized capabilities spanning visual grounding, agent environment simulation, and large-scale multimodal reasoning, enabling customers to deploy high-performance, scalable AI solutions on AWS infrastructure.

These models address different enterprise AI challenges with specialized capabilities:

LocateAnything-3B is optimized for fast, high-quality visual grounding and object localization from natural language instructions. It uses a Parallel Box Decoding (PBD) framework that decodes bounding boxes and points as atomic units in a single step, preserving geometric coherence and unlocking substantial parallelism. It enables precise object localization, dense detection, and point-based localization across diverse domains in both Enterprise Intelligence and Physical AI applications.

Qwen-AgentWorld-35B-A3B excels in simulating agent environments across seven interaction domains: tool calling, search, terminal, software engineering, Android, web, and OS interaction. It is the first language world model to cover all seven domains within a single model, predicting next environment states given an agent’s action and interaction history via long chain-of-thought reasoning—trained on over 10 million real-world interaction trajectories.

Qwen3.5-122B-A10B provides high-performance multimodal reasoning with production-friendly efficiency. It features 122B total parameters with only 10B activated per token through a hybrid architecture integrating Gated Delta Networks with sparse Mixture-of-Experts (256 experts), delivering strong reasoning, coding, agents, and visual understanding performance with a native 262K context window and minimal latency overhead.

With SageMaker JumpStart, customers can deploy any of these models with just a few clicks to address their specific AI use cases.

To get started with these models, navigate to the SageMaker JumpStart model catalog in the SageMaker console or use the SageMaker Python SDK to deploy the models to your AWS account. For more information about deploying and using foundation models in SageMaker JumpStart, see the Amazon SageMaker JumpStart documentation.

 

 

​NVIDIA’s LocateAnything-3B, Qwen’s Qwen-AgentWorld-35B-A3B, and Qwen’s Qwen3.5-122B-A10B models are now available on Amazon SageMaker JumpStart, expanding the portfolio of foundation models available to AWS customers. These three models bring specialized capabilities spanning visual grounding, agent environment simulation, and large-scale multimodal reasoning, enabling customers to deploy high-performance, scalable AI solutions on AWS infrastructure.
These models address different enterprise AI challenges with specialized capabilities:
LocateAnything-3B is optimized for fast, high-quality visual grounding and object localization from natural language instructions. It uses a Parallel Box Decoding (PBD) framework that decodes bounding boxes and points as atomic units in a single step, preserving geometric coherence and unlocking substantial parallelism. It enables precise object localization, dense detection, and point-based localization across diverse domains in both Enterprise Intelligence and Physical AI applications.
Qwen-AgentWorld-35B-A3B excels in simulating agent environments across seven interaction domains: tool calling, search, terminal, software engineering, Android, web, and OS interaction. It is the first language world model to cover all seven domains within a single model, predicting next environment states given an agent’s action and interaction history via long chain-of-thought reasoning—trained on over 10 million real-world interaction trajectories.
Qwen3.5-122B-A10B provides high-performance multimodal reasoning with production-friendly efficiency. It features 122B total parameters with only 10B activated per token through a hybrid architecture integrating Gated Delta Networks with sparse Mixture-of-Experts (256 experts), delivering strong reasoning, coding, agents, and visual understanding performance with a native 262K context window and minimal latency overhead.
With SageMaker JumpStart, customers can deploy any of these models with just a few clicks to address their specific AI use cases.
To get started with these models, navigate to the SageMaker JumpStart model catalog in the SageMaker console or use the SageMaker Python SDK to deploy the models to your AWS account. For more information about deploying and using foundation models in SageMaker JumpStart, see the Amazon SageMaker JumpStart documentation.
   

Publicado el — Deja un comentario

NVIDIA Nemotron 3.5 Lightning model is now available on Amazon SageMaker JumpStart

NVIDIA’s Nemotron 3.5 Lightning is now available on Amazon SageMaker JumpStart, giving AWS customers access to the fastest open model in its class for persistent agent workloads and rapid task execution.

Nemotron 3.5 Lightning is engineered for persistent agents and high-throughput enterprise automation across domains including personal assistants, financial document processing, cybersecurity triage, and telecom operations. Built on a hybrid Mixture-of-Experts (MoE) architecture with 30B total parameters and just 3B active per forward pass, it achieves up to 4x the throughput (~410 tokens/sec) and 30% faster task completion over comparable models. Distilled from Nemotron 3 Ultra, it handles up to 1M tokens of context via DFlash speculative decoding and integrates directly with popular agent harnesses. The model is fully open-trained on open datasets thereby allowing enterprises to post-train for their own tools, workflows, and policies, and deploy with complete ownership across edge, on-premises, or cloud infrastructure.

With SageMaker JumpStart, customers can deploy this model in a few clicks to power their specific AI workloads.

To get started with these models, navigate to the SageMaker JumpStart model catalog in the SageMaker console or use the SageMaker Python SDK to deploy the models to your AWS account. For more information about deploying and using foundation models in SageMaker JumpStart, see the Amazon SageMaker JumpStart documentation.

 

​NVIDIA’s Nemotron 3.5 Lightning is now available on Amazon SageMaker JumpStart, giving AWS customers access to the fastest open model in its class for persistent agent workloads and rapid task execution.
Nemotron 3.5 Lightning is engineered for persistent agents and high-throughput enterprise automation across domains including personal assistants, financial document processing, cybersecurity triage, and telecom operations. Built on a hybrid Mixture-of-Experts (MoE) architecture with 30B total parameters and just 3B active per forward pass, it achieves up to 4x the throughput (~410 tokens/sec) and 30% faster task completion over comparable models. Distilled from Nemotron 3 Ultra, it handles up to 1M tokens of context via DFlash speculative decoding and integrates directly with popular agent harnesses. The model is fully open-trained on open datasets thereby allowing enterprises to post-train for their own tools, workflows, and policies, and deploy with complete ownership across edge, on-premises, or cloud infrastructure.
With SageMaker JumpStart, customers can deploy this model in a few clicks to power their specific AI workloads.
To get started with these models, navigate to the SageMaker JumpStart model catalog in the SageMaker console or use the SageMaker Python SDK to deploy the models to your AWS account. For more information about deploying and using foundation models in SageMaker JumpStart, see the Amazon SageMaker JumpStart documentation.  

Publicado el — Deja un comentario

Amazon Connect Customer launches performance dashboard for Cases

Amazon Connect Customer now provides a performance dashboard for cases that helps managers monitor case volume, resolution trends, and performance against service level agreement (SLA) targets. Managers can compare current and prior-period performance across metrics such as cases created, average resolution time, first-contact resolution percentage, and SLA achievement rate. They can also analyze trends across dimensions such as case template, assigned user, or assigned queue. For example, a manager can identify that the billing team missed more SLA targets for refund cases than in the prior period, investigate the causes, and prioritize process improvements.

Cases is available in the following AWS regions: US East (N. Virginia), US West (Oregon), Canada (Central), Europe (Frankfurt), Europe (London), Asia Pacific (Seoul), Asia Pacific (Singapore), Asia Pacific (Sydney), Asia Pacific (Tokyo), and Africa (Cape Town). To learn more and get started, visit the Cases webpage and documentation.

 

​Amazon Connect Customer now provides a performance dashboard for cases that helps managers monitor case volume, resolution trends, and performance against service level agreement (SLA) targets. Managers can compare current and prior-period performance across metrics such as cases created, average resolution time, first-contact resolution percentage, and SLA achievement rate. They can also analyze trends across dimensions such as case template, assigned user, or assigned queue. For example, a manager can identify that the billing team missed more SLA targets for refund cases than in the prior period, investigate the causes, and prioritize process improvements.
Cases is available in the following AWS regions: US East (N. Virginia), US West (Oregon), Canada (Central), Europe (Frankfurt), Europe (London), Asia Pacific (Seoul), Asia Pacific (Singapore), Asia Pacific (Sydney), Asia Pacific (Tokyo), and Africa (Cape Town). To learn more and get started, visit the Cases webpage and documentation.  

Publicado el — Deja un comentario

langcache-embed-v3-small, Mellum2-12B-A2.5B-Thinking, and LightOnOCR-2-1B models now available on Amazon SageMaker JumpStart

Redis’s langcache-embed-v3-small, JetBrains’ Mellum2-12B-A2.5B-Thinking, and LightOn’s LightOnOCR-2-1B models are now available on Amazon SageMaker JumpStart, expanding the portfolio of foundation models available to AWS customers. These three models bring specialized capabilities spanning semantic caching optimization, code-focused reasoning, and end-to-end document OCR, enabling customers to deploy high-performance, scalable AI solutions on AWS infrastructure.

langcache-embed-v3-small is optimized for semantic caching in LLM applications. It maps sentences and paragraphs into a dense vector space purpose-built for identifying semantically equivalent queries regardless of phrasing, enabling intelligent cache hits that reduce redundant LLM calls and accelerate response times in high-volume inference workloads.

Mellum2-12B-A2.5B-Thinking excels in code generation, debugging, multi-step reasoning, and agentic coding workflows. It uses a Mixture-of-Experts architecture (64 experts, 8 activated per token), activating only 2.5B of its 12B total parameters per forward pass with a 131,072-token context length. It emits explicit chain-of-thought reasoning traces before final answers, delivering high-throughput, low-latency inference ideal for routing, RAG, sub-agents, and private deployments.

LightOnOCR-2-1B provides end-to-end multilingual document-to-text conversion for PDFs, scans, and images without brittle OCR pipelines. This 1B-parameter vision-language model directly transduces page images into clean, naturally ordered text, achieving state-of-the-art performance on OlmOCR-Bench while being ~9× smaller and significantly faster than competing approaches.

With SageMaker JumpStart, customers can deploy any of these models with just a few clicks to address their specific AI use cases.

To get started with these models, navigate to the SageMaker JumpStart model catalog in the SageMaker console or use the SageMaker Python SDK to deploy the models to your AWS account. For more information about deploying and using foundation models in SageMaker JumpStart, see the Amazon SageMaker JumpStart documentation.

 

​Redis’s langcache-embed-v3-small, JetBrains’ Mellum2-12B-A2.5B-Thinking, and LightOn’s LightOnOCR-2-1B models are now available on Amazon SageMaker JumpStart, expanding the portfolio of foundation models available to AWS customers. These three models bring specialized capabilities spanning semantic caching optimization, code-focused reasoning, and end-to-end document OCR, enabling customers to deploy high-performance, scalable AI solutions on AWS infrastructure.
langcache-embed-v3-small is optimized for semantic caching in LLM applications. It maps sentences and paragraphs into a dense vector space purpose-built for identifying semantically equivalent queries regardless of phrasing, enabling intelligent cache hits that reduce redundant LLM calls and accelerate response times in high-volume inference workloads.
Mellum2-12B-A2.5B-Thinking excels in code generation, debugging, multi-step reasoning, and agentic coding workflows. It uses a Mixture-of-Experts architecture (64 experts, 8 activated per token), activating only 2.5B of its 12B total parameters per forward pass with a 131,072-token context length. It emits explicit chain-of-thought reasoning traces before final answers, delivering high-throughput, low-latency inference ideal for routing, RAG, sub-agents, and private deployments.
LightOnOCR-2-1B provides end-to-end multilingual document-to-text conversion for PDFs, scans, and images without brittle OCR pipelines. This 1B-parameter vision-language model directly transduces page images into clean, naturally ordered text, achieving state-of-the-art performance on OlmOCR-Bench while being ~9× smaller and significantly faster than competing approaches.
With SageMaker JumpStart, customers can deploy any of these models with just a few clicks to address their specific AI use cases. To get started with these models, navigate to the SageMaker JumpStart model catalog in the SageMaker console or use the SageMaker Python SDK to deploy the models to your AWS account. For more information about deploying and using foundation models in SageMaker JumpStart, see the Amazon SageMaker JumpStart documentation.