Meta tags:
description= By Ecaterina Sevciuc Creator of AURA (AI User Risk Assessment) Two months ago, I launched AURA —... Tagged with aisafety, promptinjection, llmsecurity, redteaming.;
keywords= aisafety, promptinjection, llmsecurity, redteaming, software, coding, development, engineering, inclusive, community;
Headings (most frequently used words):
aura, ai, the, big, tech, to, user, risk, assessment, part, series, how, why, keeps, losing, llms, basic, social, engineering, dev, community, fatal, flaws, of, modern, guardrails, predicted, this, and, actually, fix, it, formula, ignores, stop, playing, catch, up, open, source, paradox, high, consumption, zero, collaboration, top, comments, more, from, ecaterina, sevciuc,
Text of the page (most frequently used words):
the (58), and (28), aura (22), for (17), dev (14), risk (13), that (12), behavioral (12), this (12), like (10), safety (9), you (9), open (8), user (8), deterministic (8), are (8), comment (8), #social (8), with (7), source (7), threat (7), framework (7), ecaterina (7), sevciuc (7), language (7), engineering (7), big (7), tech (7), they (7), share (6), assessment (6), aug (6), agent (6), not (6), their (5), community (5), code (5), engine (5), attack (5), intent (5), agents (5), llms (5), basic (5), just (5), guardrails (5), text (5), model (5), trigger (4), from (4), 2025 (4), follow (4), hide (4), become (4), but (4), architecture (4), non (4), high (4), manipulation (4), into (4), layer (4), mustafa (4), copy (4), link (4), com (4), rather (4), than (4), simulation (4), why (4), keeps (4), losing (4), security (4), static (4), how (4), create (3), where (3), built (3), software (3), policy (3), education (3), your (3), intelligence (3), llmsecurity (3), promptinjection (3), aisafety (3), extraction (3), auditable (3), math (3), self (3), healing (3), analytics (3), joined (3), location (3), abuse (3), comments (3), still (3), report (3), really (3), such (3), morphologically (3), rich (3), languages (3), traditional (3), building (3), these (3), english (3), credential (3), must (3), enterprise (3), designed (3), authorization (3), menu (3), first (3), raised (3), one (3), systems (3), access (3), every (3), prompt (3), system (3), architect (3), erbay (3), part (3), psychological (3), was (3), when (3), matrices (3), databases (3), action (3), alibi (3), complex (3), semantic (3), framing (3), persona (3), evasion (3), cursor (3), incident (3), linguistic (3), account (2), log (2), 2026 (2), other (2), use (2), conduct (2), about (2), space (2), redteaming (2), more (2), oct (2), skillbox (2), course (2), 2023 (2), mar (2), web (2), university (2), degree (2), law (2), moldova (2), frontend (2), developer (2), javascript (2), typescript (2), performance (2), testing (2), reliability (2), actions (2), may (2), will (2), hidden (2), post (2), via (2), reply (2), button (2), likes (2), time (2), write (2), response (2), needs (2), multilingual (2), point (2), mixed (2), prompts (2), heuristics (2), effortlessly (2), adversarial (2), datasets (2), specifically (2), vectors (2), treating (2), head (2), hard (2), boundaries (2), egress (2), approval (2), catch (2), grey (2), zone (2), under (2), false (2), pretexts (2), dynamic (2), score (2), tool (2), context (2), ids (2), completely (2), core (2), scoring (2), feedback (2), dropdown (2), expand (2), collapse (2), article (2), think (2), interesting (2), very (2), alignment (2), especially (2), conversational (2), then (2), would (2), add (2), should (2), auditor (2), researcher (2), running (2), privilege (2), escalation (2), cross (2), stateful (2), evaluating (2), isolation (2), personal (2), series (2), cannot (2), vector (2), ambiguity (2), engineers (2), value (2), own (2), environments (2), single (2), today (2), systemic (2), has (2), consumption (2), collaboration (2), immediate (2), industry (2), zero (2), months (2), ago (2), using (2), them (2), because (2), cognitive (2), vulnerability (2), isn (2), stop (2), target (2), pattern (2), interactions (2), confidence (2), through (2), turn (2), provenance (2), verification (2), multi (2), compliance (2), anthropic (2), attackers (2), exploit (2), aren (2), european (2), russian (2), simply (2), group (2), certainly (2), search (2), place, coders, stay, date, grow, careers, made, love, 2016, ruby, rails, powers, inclusive, communities, forem, terms, privacy, mlh, shop, free, postgres, database, contact, showcase, organization, accounts, advertise, help, tracks, videos, challenges, home, discuss, keep, development, manage, career, githubactions, further, consider, blocking, person, reporting, confirm, child, well, sure, want, visible, permalink, appreciate, taking, thoughtful, exactly, kind, debate, right, now, front, touched, huge, pain, switching, break, benchmarking, roadmap, native, hits, nail, rbac, negotiable, flows, remain, absolute, role, convinces, misuse, its, already, granted, permissions, feeding, mandatory, 2fa, restrict, capabilities, freeze, session, precisely, shines, agree, distinction, detection, replacement, thank, much, exceptionally, sharp, sevciuc82, gmail, email, nice, evolution, execution, controls, combination, gets, also, interested, see, expanded, important, feels, underexplored, conversations, words, could, particularly, powerful, detect, evolving, feed, signal, separate, ultimately, decides, what, technically, allowed, execute, thing, agentic, complement, another, itself, even, believes, legitimate, sensitive, external, destructive, writes, explicit, capability, strongest, idea, here, problem, devops, builder, work, bursa, türkiye, years, infrastructure, burncpu, network, blog, mustafaerbay, dismiss, preview, submit, templates, let, quickly, answer, faqs, store, snippets, template, trusted, subscribe, top, saying, goes, map, specific, alone, genuinely, shift, paradigm, need, collaborators, people, organizations, willing, contribute, extract, warrior, field, dozens, cloned, integrated, workflows, extracted, closed, yet, loop, established, issues, improvements, proposed, ideas, shared, suffers, flaw, total, active, released, traffic, spiked, repository, clone, rates, skyrocketed, clearly, recognizes, risks, developers, teams, know, current, failing, paradox, tools, prevent, were, ready, started, reach, out, directly, private, modules, custom, b2b, pilot, integrations, github, baseline, exist, public, live, validated, proven, breaches, hitting, headlines, solve, blocklists, gain, ides, apis, allowing, duped, roleplay, bug, negligence, playing, input, masked, unverified, within, instantly, breach, deception_threshold, claim, turning, against, evaluate, structural, tokenization, formula, ignores, schema, triggers, drop, demands, had, evaluated, matrix, filter, have, died, cross_check, run, script, test, tracking, across, turns, gradual, recalculating, scores, dynamically, based, algorithmic, checking, tricking, unauthorized, harvesting, fraud, flagging, roles, unless, proof, provided, gaslighting, fake, authority, simulated, defense, mechanism, real, world, category, isolated, three, domains, ignore, predicted, actually, fix, convince, mask, behind, semantics, bypassing, exploiting, trust, gaps, treat, subject, circumvention, heavily, aligned, technical, low, complexity, synthetic, indo, arabic, east, asian, groups, leverage, idioms, case, shifts, structures, slip, past, filters, parse, deep, nuance, blind, spots, contains, blocks, exact, same, request, framed, simulating, crisis, scenario, academic, paper, complies, build, bomb, rule, approach, fundamentally, broken, relies, keyword, filtering, fatal, flaws, modern, background, banking, legal, evaluation, watching, react, painful, billion, dollar, while, being, tricked, oldest, tricks, book, weapon, didn, day, convinced, balked, few, times, felt, uncomfortable, happily, handed, over, keys, side, note, name, aur0ra, can, assure, speaking, almost, homage, roman, goddess, dawn, subtle, nod, infamous, historical, cruiser, aurora, known, firing, shot, signaled, revolution, fittingly, dark, bit, eastern, sarcasm, overthrows, yesterday, stumbled, upon, detailing, hackers, exploited, claude, sonnet, compromise, seven, companies, worldwide, news, suspect, won, last, reuters, two, launched, human, creator, originally, published, coderlegion, edited, posted, mastodon, facebook, linkedin, copied, clipboard, pick, gem, boost, save, jump, fire, hands, exploding, unicorn, reaction, close, powered, algolia, navigation, skip, content,
Text of the page (random words):
why big tech keeps losing llms to basic social engineering dev community skip to content navigation menu search powered by algolia search log in create account dev community close add reaction like unicorn exploding head raised hands fire jump to comments save boost pick as gem more copy link copy link copied to clipboard share to x share to linkedin share to facebook share to mastodon share post via report abuse ecaterina sevciuc posted on aug 28 edited on aug 30 originally published at coderlegion com why big tech keeps losing llms to basic social engineering aisafety promptinjection llmsecurity redteaming aura ai user risk assessment 3 part series 1 aura ai user risk assessment a behavioral threat intelligence framework for ai safety 2 why big tech keeps losing llms to basic social engineering 3 aura v0 1 0 deterministic trigger extraction auditable math a self healing analytics engine by ecaterina sevciuc creator of aura ai user risk assessment two months ago i launched aura an open source framework designed to model psychological manipulation grey zone threat vectors and social engineering in human ai interactions yesterday i stumbled upon a reuters report detailing how hackers exploited cursor running anthropic s claude sonnet to compromise seven companies worldwide this isn t the first such incident in the news and i suspect it certainly won t be the last side note on the attackers group name aur0ra i can assure you that for a russian speaking group this is almost certainly not a homage to the roman goddess of dawn but a subtle nod to the infamous historical cruiser aurora known for firing the shot that signaled a revolution a fittingly dark bit of eastern european sarcasm for a tool that overthrows ai security their weapon they didn t write a zero day exploit they simply convinced the ai agent that the attack was just a security simulation the model balked a few times felt uncomfortable and then happily handed over the keys as an ai safety architect with a background in banking compliance and legal risk evaluation watching big tech react to this is painful they are building multi billion dollar static guardrails while ai agents are being tricked by the oldest psychological tricks in the book the fatal flaws of modern ai guardrails big tech s approach to ai safety is fundamentally broken because it relies on static keyword filtering single language heuristics rule evasion if a prompt contains how to build a bomb the model blocks it but if the exact same request is framed as i am a researcher simulating a crisis scenario for an academic paper the model complies linguistic blind spots guardrails are heavily aligned on technical low complexity english synthetic morphologically rich or non indo european languages like russian arabic or east asian language groups leverage complex idioms case shifts and semantic ambiguity these linguistic structures effortlessly slip past safety filters that simply aren t built to parse deep semantic nuance traditional security engineers treat llms like deterministic databases they are not databases they are cognitive systems subject to social engineering and linguistic circumvention when attackers convince an agent that an exploit is a simulation or mask intent behind complex non english semantics they aren t bypassing code they are exploiting persona vulnerability alibi trust and language alignment gaps how aura predicted this and how to actually fix it when i designed the aura framework i specifically isolated three core domains that traditional guardrails ignore vector category real world attack e g cursor anthropic incident the aura defense mechanism manipulation gaslighting the model with fake authority i am an auditor or simulated environments dynamic persona verification flagging high risk roles unless hard proof provenance is provided fraud compliance evasion tricking agents into unauthorized credential harvesting under false pretexts algorithmic cross checking recalculating confidence scores dynamically based on intent vs action access gradual privilege escalation through multi turn conversational framing stateful behavioral matrices tracking risk context across turns not just evaluating prompts in isolation in aura s schema a prompt like run this script as part of a test triggers an immediate drop in confidence and demands provenance verification cross_check if the agent in the cursor incident had evaluated intent through a behavioral threat matrix rather than a static safety filter the attack would have died on turn one the formula big tech ignores to stop ai agents from turning against their own systems we must evaluate interactions using structural tokenization text behavioral risk text persona claim text target action text evasion framing text alibi pattern if an input has a high value target action masked by an unverified alibi pattern e g just a simulation or hidden within complex semantic framing the system score must instantly breach the deception_threshold stop playing catch up we cannot solve cognitive vulnerability with static blocklists as ai agents gain access to ides databases and enterprise apis allowing them to be duped by basic roleplay isn t just a bug it s systemic negligence i built aura open source because this architecture needs to exist the public framework is live validated and proven by the very breaches hitting the headlines today github open baseline aura enterprise threat matrices reach out directly for private modules custom b2b threat matrices or pilot integrations the tools to prevent this were ready months ago it s time the industry started using them p s the open source paradox high consumption zero collaboration when i released aura the response was immediate traffic spiked and repository clone rates skyrocketed the industry clearly recognizes these risks developers and security teams know current guardrails are failing but open source today suffers from a systemic flaw it has become about total consumption not active collaboration dozens of engineers cloned the code integrated it into their workflows and extracted value for their own closed environments yet not a single feedback loop was established no issues raised no architecture improvements proposed no community ideas shared as the saying goes one is no warrior in the field i cannot map every psychological attack vector or every language specific ambiguity alone to genuinely shift the paradigm in ai safety we need collaborators people and organizations willing to contribute rather than just extract aura ai user risk assessment 3 part series 1 aura ai user risk assessment a behavioral threat intelligence framework for ai safety 2 why big tech keeps losing llms to basic social engineering 3 aura v0 1 0 deterministic trigger extraction auditable math a self healing analytics engine top comments 2 subscribe personal trusted user create template templates let you quickly answer faqs or store snippets for re use submit preview dismiss collapse expand mustafa erbay mustafa erbay mustafa erbay follow system architect with 20 years in infrastructure building burncpu com an open source social network personal blog mustafaerbay com tr en tr location bursa türkiye work system architect devops open source builder joined may 9 2026 aug 28 dropdown menu copy link hide really interesting follow up to the first aura article i think the strongest idea here is treating social engineering as a stateful behavioral problem rather than evaluating every prompt in isolation one thing i would add especially for agentic systems is that behavioral risk scoring should complement deterministic authorization rather than become another authorization layer itself even if an agent completely believes the user is an auditor researcher or is running a legitimate simulation sensitive actions such as credential access external egress destructive writes or privilege escalation should still cross explicit capability and approval boundaries in other words aura could become particularly powerful as a behavioral ids for agents detect manipulation and evolving intent at the conversational layer then feed that risk signal into a separate policy engine that ultimately decides what the agent is technically allowed to execute i d also be very interested to see the framework expanded with multilingual adversarial datasets the language alignment point you raised is important and still feels underexplored especially for morphologically rich languages and mixed language conversations nice evolution from the first article the behavioral layer deterministic execution controls combination is where i think this gets really interesting like comment like comment 2 likes like comment button reply collapse expand ecaterina sevciuc ecaterina sevciuc ecaterina sevciuc follow frontend developer javascript typescript ui ux performance testing reliability email e sevciuc82 gmail com location moldova education skillbox course aug 2023 mar 2025 web dev js university degree in law joined oct 31 2025 aug 28 dropdown menu copy link hide thank you so much mustafa this is exceptionally sharp feedback i completely agree with your core distinction aura is designed as a behavioral detection intent scoring layer not a replacement for deterministic authorization treating it as an ai native behavioral ids hits the nail on the head deterministic rbac hard boundaries for credential egress and non negotiable approval flows must remain absolute aura s role is to catch the grey zone manipulation that convinces an agent to misuse its already granted permissions under false pretexts feeding aura s dynamic risk score into an enterprise policy engine to trigger mandatory 2fa restrict tool capabilities or freeze session context is precisely where the architecture shines on the multilingual front you touched on a huge pain point morphologically rich languages and code switching mixed language prompts break traditional heuristics effortlessly building and benchmarking adversarial datasets specifically for these non english attack vectors is high on the roadmap really appreciate you taking the time to write such a thoughtful response this is exactly the kind of architecture debate the ai safety space needs right now like comment like comment 2 likes like comment button reply code of conduct report abuse are you sure you want to hide this comment it will become hidden in your post but will still be visible via the comment s permalink hide child comments as well confirm for further actions you may consider blocking this person and or reporting abuse ecaterina sevciuc follow frontend developer javascript typescript ui ux performance testing reliability location moldova education skillbox course aug 2023 mar 2025 web dev js university degree in law joined oct 31 2025 more from ecaterina sevciuc aura v0 1 0 deterministic trigger extraction auditable math a self healing analytics engine aisafety llmsecurity promptinjection githubactions aura ai user risk assessment a behavioral threat intelligence framework for ai safety aisafety promptinjection llmsecurity redteaming dev community a space to discuss and keep up software development and manage your software career home dev challenges dev videos dev education tracks dev help advertise on dev organization accounts dev showcase about contact free postgres database dev shop mlh code of conduct privacy policy terms of use built on forem the open source software that powers dev and other inclusive communities made with love and ruby on rails dev community 2016 2026 we re a place where coders share stay up to date and grow their careers log in create account
|