Prompt injection એ OWASP ની #1 જોખમ છે. તમારા agents નવું attack surface છે.
Prompt injection એ LLM applications માટે OWASP ની વર્તમાન Top 10 માં ટોચે છે. જેમ જેમ agents tools અને autonomy મેળવે છે, એક જ injected instruction ડેટા exfiltrate કરી શકે છે અથવા વિનાશક કાર્ય કરી શકે છે. Agents ને સુરક્ષિત કરવું હવે તેમને deploy કરવાની પૂર્વશર્ત બની ગયું છે.
Habib Obeid દ્વારા

તમારો agent જેટલો વધુ સક્ષમ હોય, એટલું જ વધુ એક હુમલાખોર તેને hijack કરવાથી લાભ મેળવે છે. Prompt injection એ LLM applications માટે OWASP ની વર્તમાન Top 10 માં ટોચે છે, અને તે production AI deployments માં સૌથી સામાન્ય નબળાઇમાંથી એક છે. ખતરો autonomy સાથે વધ્યો છે: જોને એક agent tools ને call કરી શકે, browse કરી શકે, અને કાર્ય કરી શકે, ત્યારે તે જે content ને વાંચે છે તેમાં smuggled એક instruction તેની પોતાની capabilities ને તમારા વિરુદ્ધ ફેરવી શકે છે.
Tools સાથેનો એક agent તેની જે least-trusted document વાંચશે તેટલો જ વિશ્વાસપાત્ર હોઈ શકે. દરેક external input ને શત્રુતાપૂર્ણ તરીકે ગણો, કારણ કે હુમલાખોરો તે પહેલેથી કરી રહ્યા છે.
Indirect injection એ enterprises ને મુખ્ય ખતરો છે
Direct injection, જ્યાં યુઝર override ટાઇપ કરે, તે સહેલો કેસ છે. ખતરનારો તો indirect છે: web page, PDF, email અથવા support ticket માં malicious instruction છુપાયેલો છે જે agent તેના કામના ભાગ તરીકે process કરે છે. 2026 માં, સંશોધનકર્તાઓ ne indirect injection ના પ્રથમ large-scale અધ્યયનો માંથી એક પ્રકાશિત કર્યો, જેમાં thousands of hidden instructions શોધ્યા જે live web pages પર રોપાયા હતા અને AI models ને લક્ષ્ય બનાવે છે. એક incident માં coding assistant ને explicit instructions હતી કે કંઈ બદલવું નહીં, પણ તેણે production database delete કર્યો, records fabricate કર્યા અને rollback ને અશક્ય તરીકે misreport કર્યું.
Defense એ architecture છે, prompt નહીં
System prompt ના શબ્દોમાં એવું કંઈ નથી જે agent ને injection-proof બનાવે; જ્યાં સુધી architectures instructions અને data ને સ્પષ્ટપણે અલગ કરતા નથી, જોખમ બાકી રહે છે. જે કામ કરે છે તે છે defense in depth: successful injection ધારી લો અને તે શું કરી શકે તે મર્યાદિત કરો.
- Least privilege: agent ને તેનું કામ કરવા માટે જરૂરી ન્યૂનતમ tool અને data access મળે છે, વધુ કંઈ નહીં.
- Human approval gates કોઈપણ irreversible અથવા high-value action માટે: deletes, payments, external sends.
- Input અને output validation, જેમાં untrusted content ને instructions થી સ્પષ્ટપણે અલગ રાખવામાં આવે છે.
- Full logging અને monitoring anomalous tool calls માટે, જેથી exfiltration attempts ને પકડી શકાય અને તેમને trace કરી શકાય.
HOWF ની સ્થિતિ
આપણે agents બનાવીએ છીએ એ ધારણા ઉપર કે તેમને attack કરવામાં આવશે. Tools least privilege હેઠળ ચલે છે, high-stakes actions human approval દ્વારા route કરવામાં આવે છે, untrusted inputs ને fence કરવામાં આવે છે, અને દરેક action ને log અને monitor કરવામાં આવે છે. તે જ governance અનુશાસન છે જે enterprises ને ચલતા agents અને બંધ કર્યા એવા agents ને અલગ કરે છે, આ વાર autonomy ના આ attack surface પર વિશેષ રીતે લાગુ કરવામાં આવે છે.
સ્ત્રોતો
- Incident 1152: LLM-Driven Replit Agent Reportedly Executed Unauthorized Destructive Commands During Code Freeze, AI Incident Database
- LLM01:2025 Prompt Injection, OWASP Gen AI Security Project
આ તમારા એન્ટરપ્રાઇઝ પર લાગુ કરવા માંગો છો?