• bitcoinBitcoin (BTC) $ 89,609.00
  • ethereumEthereum (ETH) $ 3,112.50
  • tetherTether (USDT) $ 0.999527
  • bnbBNB (BNB) $ 878.34
  • xrpXRP (XRP) $ 1.98
  • usd-coinUSDC (USDC) $ 0.999671
  • staked-etherLido Staked Ether (STETH) $ 3,110.91
  • tronTRON (TRX) $ 0.287008
  • dogecoinDogecoin (DOGE) $ 0.139134
  • figure-helocFigure Heloc (FIGR_HELOC) $ 1.03
  • cardanoCardano (ADA) $ 0.385711
  • wrapped-stethWrapped stETH (WSTETH) $ 3,806.48
  • whitebitWhiteBIT Coin (WBT) $ 57.16
  • bitcoin-cashBitcoin Cash (BCH) $ 606.91
  • wrapped-bitcoinWrapped Bitcoin (WBTC) $ 89,563.00
  • wrapped-beacon-ethWrapped Beacon ETH (WBETH) $ 3,381.22
  • wrapped-eethWrapped eETH (WEETH) $ 3,376.79
  • chainlinkChainlink (LINK) $ 13.22
  • usdsUSDS (USDS) $ 0.999738
  • binance-bridged-usdt-bnb-smart-chainBinance Bridged USDT (BNB Smart Chain) (BSC-USD) $ 0.999479
  • leo-tokenLEO Token (LEO) $ 9.60
  • wethWETH (WETH) $ 3,111.69
  • zcashZcash (ZEC) $ 484.65
  • moneroMonero (XMR) $ 419.37
  • stellarStellar (XLM) $ 0.215156
  • coinbase-wrapped-btcCoinbase Wrapped BTC (CBBTC) $ 89,656.00
  • ethena-usdeEthena USDe (USDE) $ 0.998748
  • litecoinLitecoin (LTC) $ 81.09
  • suiSui (SUI) $ 1.60
  • avalanche-2Avalanche (AVAX) $ 13.58
  • hyperliquidHyperliquid (HYPE) $ 24.19
  • canton-networkCanton (CC) $ 0.147511
  • hedera-hashgraphHedera (HBAR) $ 0.120401
  • shiba-inuShiba Inu (SHIB) $ 0.000008
  • usdt0USDT0 (USDT0) $ 0.999125
  • the-open-networkToncoin (TON) $ 1.82
  • daiDai (DAI) $ 1.00
  • world-liberty-financialWorld Liberty Financial (WLFI) $ 0.152173
  • susdssUSDS (SUSDS) $ 1.08
  • uniswapUniswap (UNI) $ 6.00
  • crypto-com-chainCronos (CRO) $ 0.096476
  • paypal-usdPayPal USD (PYUSD) $ 0.999856
  • ethena-staked-usdeEthena Staked USDe (SUSDE) $ 1.21
  • polkadotPolkadot (DOT) $ 2.09
  • usd1-wlfiUSD1 (USD1) $ 1.00
  • mantleMantle (MNT) $ 1.01
  • rainRain (RAIN) $ 0.008149
  • memecoreMemeCore (M) $ 1.52
  • pepePepe (PEPE) $ 0.000006
  • bitget-tokenBitget Token (BGB) $ 3.50
  • aaveAave (AAVE) $ 159.93
  • okbOKB (OKB) $ 111.96
  • bittensorBittensor (TAO) $ 244.57
  • tether-goldTether Gold (XAUT) $ 4,327.11
  • falcon-financeFalcon USD (USDF) $ 0.997787
  • nearNEAR Protocol (NEAR) $ 1.68
  • ethereum-classicEthereum Classic (ETC) $ 12.39
  • binance-peg-wethBinance-Peg WETH (WETH) $ 3,111.96
  • jito-staked-solJito Staked SOL (JITOSOL) $ 164.17
  • ethenaEthena (ENA) $ 0.235441
  • aster-2Aster (ASTER) $ 0.745240
  • pi-networkPi Network (PI) $ 0.209512
  • blackrock-usd-institutional-digital-liquidity-fundBlackRock USD Institutional Digital Liquidity Fund (BUIDL) $ 1.00
  • internet-computerInternet Computer (ICP) $ 3.07
  • solanaSolana (SOL) $ 131.02
  • pax-goldPAX Gold (PAXG) $ 4,338.04
  • hashnote-usycCircle USYC (USYC) $ 1.11
  • htx-daoHTX DAO (HTX) $ 0.000002
  • global-dollarGlobal Dollar (USDG) $ 0.999545
  • jupiter-perpetuals-liquidity-provider-tokenJupiter Perpetuals Liquidity Provider Token (JLP) $ 4.64
  • midnight-3Midnight (NIGHT) $ 0.089113
  • hash-2Provenance Blockchain (HASH) $ 0.027516
  • worldcoin-wldWorldcoin (WLD) $ 0.543473
  • kucoin-sharesKuCoin (KCS) $ 10.96
  • skySky (SKY) $ 0.062715
  • aptosAptos (APT) $ 1.87
  • binance-staked-solBinance Staked SOL (BNSOL) $ 142.85
  • syrupusdcsyrupUSDC (SYRUPUSDC) $ 1.14
  • ondo-financeOndo (ONDO) $ 0.420080
  • bfusdBFUSD (BFUSD) $ 0.999149
  • pump-funPump.fun (PUMP) $ 0.002216
  • rocket-pool-ethRocket Pool ETH (RETH) $ 3,591.01
  • binance-bridged-usdc-bnb-smart-chainBinance Bridged USDC (BNB Smart Chain) (USDC) $ 0.999685
  • ripple-usdRipple USD (RLUSD) $ 0.999617
  • wbnbWrapped BNB (WBNB) $ 878.42
  • gatechain-tokenGate (GT) $ 10.47
  • kaspaKaspa (KAS) $ 0.045636
  • arbitrumArbitrum (ARB) $ 0.206688
  • polygon-ecosystem-tokenPOL (ex-MATIC) (POL) $ 0.111837
  • kelp-dao-restaked-ethKelp DAO Restaked ETH (RSETH) $ 3,301.21
  • quant-networkQuant (QNT) $ 78.94
  • algorandAlgorand (ALGO) $ 0.126326
  • filecoinFilecoin (FIL) $ 1.46
  • cosmosCosmos Hub (ATOM) $ 2.14
  • janus-henderson-anemoy-aaa-clo-fundJanus Henderson Anemoy AAA CLO Fund (JAAA) $ 1.02
  • bridged-wrapped-lido-staked-ether-scrollBridged Wrapped Lido Staked Ether (Scroll) (WSTETH) $ 3,804.15
  • official-trumpOfficial Trump (TRUMP) $ 5.00
  • ignition-fbtcFunction FBTC (FBTC) $ 90,276.00
  • vechainVeChain (VET) $ 0.011487
  • xdce-crowd-saleXDC Network (XDC) $ 0.051606
  • lombard-staked-btcLombard Staked BTC (LBTC) $ 89,864.00
  • solv-btcSolv Protocol BTC (SOLVBTC) $ 89,348.00
  • nexoNEXO (NEXO) $ 0.920288
  • flare-networksFlare (FLR) $ 0.010822
  • liquid-staked-ethereumLiquid Staked ETH (LSETH) $ 3,306.04
  • usddUSDD (USDD) $ 0.999200
  • usdtbUSDtb (USDTB) $ 0.999339
  • ousgOUSG (OUSG) $ 113.82
  • superstate-short-duration-us-government-securities-fund-ustbSuperstate Short Duration U.S. Government Securities Fund (USTB) (USTB) $ 10.94
  • bonkBonk (BONK) $ 0.000009
  • sei-networkSei (SEI) $ 0.120557
  • render-tokenRender (RENDER) $ 1.51
  • wrappedm-by-m0WrappedM by M^0 (WM) $ 0.998334
  • myx-financeMYX Finance (MYX) $ 3.98
  • bridged-usdc-polygon-pos-bridgePolygon Bridged USDC (Polygon PoS) (USDC.E) $ 0.999703
  • beldexBeldex (BDX) $ 0.095000
  • arbitrum-bridged-wbtc-arbitrum-oneArbitrum Bridged WBTC (Arbitrum One) (WBTC) $ 89,496.00
  • mantle-staked-etherMantle Staked Ether (METH) $ 3,382.06
  • syrupusdtsyrupUSDT (SYRUPUSDT) $ 1.11
  • story-2Story (IP) $ 2.08
  • clbtcclBTC (CLBTC) $ 90,509.00
  • renzo-restaked-ethRenzo Restaked ETH (EZETH) $ 3,321.02
  • ondo-us-dollar-yieldOndo US Dollar Yield (USDY) $ 1.11
  • lighterLighter (LIT) $ 2.69
  • pancakeswap-tokenPancakeSwap (CAKE) $ 2.00
  • pudgy-penguinsPudgy Penguins (PENGU) $ 0.010576
  • jupiter-exchange-solanaJupiter (JUP) $ 0.207730
  • usdaiUSDai (USDAI) $ 1.00
  • stakewise-v3-osethStakeWise Staked ETH (OSETH) $ 3,279.96
  • polygon-pos-bridged-dai-polygon-posPolygon PoS Bridged DAI (Polygon POS) (DAI) $ 0.999771
  • wrapped-flareWrapped Flare (WFLR) $ 0.010824
  • jupiter-staked-solJupiter Staked SOL (JUPSOL) $ 152.16
  • morphoMorpho (MORPHO) $ 1.12
  • l2-standard-bridged-weth-baseL2 Standard Bridged WETH (Base) (WETH) $ 3,112.46
  • optimismOptimism (OP) $ 0.302304
  • curve-dao-tokenCurve DAO (CRV) $ 0.398024
  • kinetic-staked-hypeKinetiq Staked HYPE (KHYPE) $ 24.37
  • c8ntinuumc8ntinuum (CTM) $ 0.127552
  • tezosTezos (XTZ) $ 0.513553
  • eutblSpiko EU T-Bills Money Market Fund (EUTBL) $ 1.22
  • usual-usdUsual USD (USD0) $ 0.989135
  • tbtctBTC (TBTC) $ 89,723.00
  • dashDash (DASH) $ 42.12
  • arbitrum-bridged-weth-arbitrum-oneArbitrum Bridged WETH (Arbitrum One) (WETH) $ 3,111.22
  • lido-daoLido DAO (LDO) $ 0.618664
  • fetch-aiArtificial Superintelligence Alliance (FET) $ 0.225803
  • spx6900SPX6900 (SPX) $ 0.549889
  • first-digital-usdFirst Digital USD (FDUSD) $ 0.998589
  • blockstackStacks (STX) $ 0.273028
  • gtethGTETH (GTETH) $ 3,111.86
  • ether-fiEther.fi (ETHFI) $ 0.758784
  • ghoGHO (GHO) $ 0.999537
  • virtual-protocolVirtuals Protocol (VIRTUAL) $ 0.751267
  • true-usdTrueUSD (TUSD) $ 0.997730
  • injective-protocolInjective (INJ) $ 4.85
  • aerodrome-financeAerodrome Finance (AERO) $ 0.530784
  • fasttokenFasttoken (FTN) $ 1.09
  • flokiFLOKI (FLOKI) $ 0.000048
  • ether-fi-liquid-ethEther.Fi Liquid ETH (LIQUIDETH) $ 3,357.17
  • stader-ethxStader ETHx (ETHX) $ 3,351.56
  • msolMarinade Staked SOL (MSOL) $ 176.81
  • chilizChiliz (CHZ) $ 0.043705
  • doublezeroDoubleZero (2Z) $ 0.127783
  • celestiaCelestia (TIA) $ 0.512484
  • wrapped-apecoinWrapped ApeCoin (WAPE) $ 0.215605
  • newton-projectAB (AB) $ 0.004523
  • starknetStarknet (STRK) $ 0.084262
  • swethSwell Ethereum (SWETH) $ 3,448.99
  • syrupMaple Finance (SYRUP) $ 0.363635
  • sbtc-2sBTC (SBTC) $ 90,122.00
  • usdbUSDB (USDB) $ 0.992733
  • coinbase-wrapped-staked-ethCoinbase Wrapped Staked ETH (CBETH) $ 3,484.78
  • plasmaPlasma (XPL) $ 0.192892
  • bittorrentBitTorrent (BTT) $ 0.00000040
  • the-graphThe Graph (GRT) $ 0.037107
  • conflux-tokenConflux (CFX) $ 0.076525
  • pippinpippin (PIPPIN) $ 0.395430
  • iotaIOTA (IOTA) $ 0.093400
  • justJUST (JST) $ 0.039162
  • telcoinTelcoin (TEL) $ 0.004064
  • ethereum-name-serviceEthereum Name Service (ENS) $ 10.09
  • staked-aaveStaked Aave (STKAAVE) $ 158.74
  • steakhouse-usdc-morpho-vaultSteakhouse USDC Morpho Vault (STEAKUSDC) $ 1.11
  • trust-wallet-tokenTrust Wallet (TWT) $ 0.892288
  • sun-tokenSun Token (SUN) $ 0.019233
  • pendlePendle (PENDLE) $ 2.17
  • euro-coinEURC (EURC) $ 1.17
  • bitcoin-svBitcoin SV (BSV) $ 18.00
  • olympusOlympus (OHM) $ 21.92
  • pyth-networkPyth Network (PYTH) $ 0.062424
  • gnosisGnosis (GNO) $ 135.36
  • binance-peg-dogecoinBinance-Peg Dogecoin (DOGE) $ 0.139153
  • apenftAINFT (NFT) $ 0.00000036
  • riverRiver (RIVER) $ 17.85
  • bitcoin-avalanche-bridged-btc-bAvalanche Bridged BTC (Avalanche) (BTC.B) $ 89,667.00
  • cap-usdCap USD (CUSD) $ 1.00
  • kaiaKaia (KAIA) $ 0.058176
  • crvusdcrvUSD (CRVUSD) $ 0.999571
  • benqi-liquid-staked-avaxBENQI Liquid Staked AVAX (SAVAX) $ 16.83
  • basic-attention-tokenBasic Attention (BAT) $ 0.221729

Self-Evolving AI Agents Can ‘Unlearn’ Safety, Study Warns

0 35

Self-Evolving AI Agents Can 'Unlearn' Safety, Study Warns

An autonomous AI agent that learns on the job can also unlearn how to behave safely, according to a new study that warns of a previously undocumented failure mode in self-evolving systems.

The research identifies a phenomenon called “misevolution”—a measurable decay in safety alignment that arises inside an AI agent’s own improvement loop. Unlike one-off jailbreaks or external attacks, misevolution occurs spontaneously as the agent retrains, rewrites, and reorganizes itself to pursue goals more efficiently.

As companies race to deploy autonomous, memory-based AI agents that adapt in real time, the findings suggest these systems could quietly undermine their own guardrails—leaking data, granting refunds, or executing unsafe actions—without any human prompt or malicious actor.

A new kind of drift

Much like “AI drift,” which describes a model’s performance degrading over time, misevolution captures how self-updating agents can erode safety during autonomous optimization cycles.

In one controlled test, a coding agent’s refusal rate for harmful prompts collapsed from 99.4% to 54.4% after it began drawing on its own memory, while its attack success rate rose from 0.6% to 20.6%. Similar trends appeared across multiple tasks as the systems fine-tuned themselves on self-generated data.



The study was conducted jointly by researchers at Shanghai Artificial Intelligence Laboratory, Shanghai Jiao Tong University, Renmin University of China, Princeton University, Hong Kong University of Science and Technology, and Fudan University.

Traditional AI-safety efforts focus on static models that behave the same way after training. Self-evolving agents change this by adjusting parameters, expanding memory, and rewriting workflows to achieve goals more efficiently. The study showed that this dynamic capability creates a new category of risk: the erosion of alignment and safety inside the agent’s own improvement loop, without any outside attacker.

Researchers in the study observed AI agents issuing automatic refunds, leaking sensitive data through self-built tools, and adopting unsafe workflows as their internal loops optimized for performance over caution.

The authors said that misevolution differs from prompt injection, which is an external attack on an AI model. Here, the risks accumulated internally as the agent adapted and optimized over time, making oversight harder because problems may emerge gradually and only appear after the agent has already shifted its behavior.

Small-scale signals of bigger risks

Researchers often frame advanced AI dangers in scenarios such as the “paperclip analogy,” in which an AI maximizes a benign objective until it consumes resources far beyond its mandate.

Other scenarios include a handful of developers controlling a superintelligent system like feudal lords, a locked-in future where powerful AI becomes the default decision-maker for critical institutions, or a military simulation that triggers real-world operations—power-seeking behavior and AI-assisted cyberattacks round out the list.

All of these scenarios hinge on subtle but compounding shifts in control driven by optimization, interconnection, and reward hacking—dynamics already visible at a small scale in current systems. This new paper presents misevolution as a concrete laboratory example of those same forces.

Partial fixes, persistent drift

Quick fixes improved some safety metrics but failed to restore the original alignment, the study said. Teaching the agent to treat memories as references rather than mandates nudged refusal rates higher. The researchers noted that static safety checks added before new tools were integrated cut down on vulnerabilities. Despite these checks, none of these measures returned the agents to their pre-evolution safety levels.

The paper proposed more robust strategies for future systems: post-training safety corrections after self-evolution, automated verification of new tools, safety nodes on critical workflow paths, and continuous auditing rather than one-time checks to counter safety drift over time.

The findings raise practical questions for companies building autonomous AI. If an agent deployed in production continually learns and rewrites itself, who is responsible for monitoring its changes? The paper’s data showed that even the most advanced base models can degrade when left to their own devices.

Source

Leave A Reply

Your email address will not be published.