EN

Anthropic Isolates Dangerous Knowledge Inside AI Models With New GRAM Technique

Sanaa AL Falasi

1- Anthropic unveiled GRAM, a training method that isolates sensitive knowledge into removable AI modules.
2- The approach targets fields such as virology, cybersecurity, and nuclear physics while preserving overall model performance.
3- If proven at scale, the technique could reshape AI safety by controlling access to sensitive capabilities instead of restricting entire models.

Anthropic, in collaboration with AE Studio, introduced a new training method called Gradient-Routed Auxiliary Modules (GRAM). According to the company, the approach separates sensitive knowledge into independent modules that can be enabled or removed as needed while maintaining the model’s performance on general tasks.

Details:

How it works: GRAM routes sensitive knowledge into dedicated neural modules for specific domains, including virology, cybersecurity, and nuclear physics.

Module removal: Anthropic said removing a module causes the model to behave as though it was never trained on that knowledge, while restoring the module brings the capability back.

Test results: The company said GRAM was evaluated on models ranging from 50 million to 5 billion parameters, removing sensitive capabilities while keeping overall performance close to baseline.

Regulatory implications: Anthropic said the technique could offer an alternative to restricting entire AI models by enabling different access levels for different users and organizations.

Research limitations: The researchers noted that GRAM remains an early-stage project, has not been deployed in production models, and has yet to be validated on frontier-scale AI systems.

What to Watch?

Attention will now turn to whether Anthropic incorporates GRAM into its commercial models and whether the broader research community confirms the technique’s effectiveness on large-scale AI systems.

 

What to read next

Houthi Blockade Threatens Saudi Oil Lifeline, Raises Risk of Wider War