The people paid to break into systems before attackers do can’t get AI to help them write a basic exploit. Threat actors on dark web forums are sharing uncensored models with zero refusals. According to SlashNext’s 2023 State of Phishing report, AI generated phishing attacks grew 4,151% in the 18 months after ChatGPT launched. The good guys are fighting with one hand tied behind their back.
Why This Is Blowing Up Now
Something shifted in late 2025. As AI models got better at writing attack code, the content filters got tighter in response. The providers made a business call: one viral screenshot of their model writing malware is worse for their valuation than a thousand researchers getting slowed down. I understand the logic. I think it’s costing us more than anyone is admitting.
The problem is specific and practical. Ask a general AI model to write a SQL injection payload for a penetration test report and most refuse. Ask it to analyze a suspicious binary and it hedges. Ask it to simulate a phishing email for security awareness training and it lectures you about ethics. These are core deliverables in legitimate offensive security work, not theoretical edge cases.
According to the ISC2 2024 Cybersecurity Workforce Study, the global security talent gap sits at 4.8 million unfilled positions. We need every tool available to close that gap. The tools are now fighting the researchers who need them most.
The Guardrails Don’t Stop Attackers. They Stop Researchers.
This is the part every CISO should pay attention to. A Fortune 500 company’s internal red team has access to enterprise model deployments with custom safety profiles negotiated into the contract. A solo bug bounty hunter or a small pen testing shop does not. The guardrails have created a class divide inside the security profession.
According to HackerOne, over one million researchers are registered across bug bounty platforms globally. Most work independently or at small shops. These are the people finding vulnerabilities before attackers do. They’re also the people most affected by AI refusals because they don’t have enterprise model agreements or legal teams to negotiate custom terms.
Compare that to the threat actor side. WormGPT and its successors have circulated on criminal forums since 2023. Uncensored models with no content filters are freely shared. The attackers are not filing support tickets or reading acceptable use policies. They adapted in 2023. Defenders are still navigating refusal messages in 2026.
I’ve seen this same dynamic in finance. Regulations meant to protect retail investors often handicap individual traders while institutions have compliance teams working around the rules before lunch. The regulation protects the brand of the regulator, not the market participant. AI guardrails are a business decision dressed up as an ethics decision, and the people paying the price are the ones doing legitimate work.
According to a 2024 report from Cybersecurity Ventures, global cybercrime costs are projected to hit $10.5 trillion annually by 2025. That’s the number on the attacker side of the ledger. The researchers trying to reduce that number are being slowed down by the tools that were supposed to help them. That math doesn’t work.
Good operators understand what’s happening here. When you make a tool less useful for defenders while attackers find workarounds in 48 hours, you haven’t made anything safer. You’ve added friction to the wrong side of the equation.
What I Would Do Right Now If I Were in Security
First, build your documentation practice. The researchers winning despite the guardrails are producing detailed writeups, proof of concept videos, and public CVE disclosures. That paper trail is what gets you enterprise model access, expanded agreements, and credibility with vendors. If you’re not already creating video walkthroughs of your findings, InVideo AI lets you turn a screen recording and notes into a professional explainer fast, without a production budget. A documented public track record changes how AI providers respond to your access requests.
Second, know which models actually support security professionals. Some providers have explicit carve-outs for security research in their usage policies. Anthropic and OpenAI both have provisions for penetration testing and research contexts. The key is precision. A vague request triggers a refusal. A specific, context-rich request with your role and scope stated clearly gets a different response most of the time.
Third, look at purpose built security AI tools. General consumer models are not built for offensive security work. Specialized tools are coming to market that are. AppSumo runs lifetime deals on security and developer tools regularly. Checking there before buying annual subscriptions on new research software has saved me real money on tools I actually use.
Fourth, engage directly with AI companies’ trust and safety teams. This sounds like slow advice but it’s not. The researchers who have shaped how these tools handle security requests are the ones who showed up, documented use cases clearly, and made the case in writing. The policies are not fixed. They move based on who’s in the room making the argument.
The Bottom Line
Attackers adapted to AI in 2023. Security researchers are still filing complaints in 2026. The guardrails are not protecting anyone from threat actors; those actors found their tools two years ago. The guardrails are protecting AI companies from bad press while independent researchers fall further behind every month. If you’re in security and you’re not building public credibility and direct relationships with AI providers, you’re already behind and the gap is widening.
Frequently Asked Questions
Are AI guardrails getting stricter or looser for security researchers?
The trend through 2025 and into 2026 has been stricter public-facing models paired with more detailed enterprise agreements. Major providers have added formal channels for security researchers to request expanded access. The public models are more restricted now; the path to useful access exists but requires documentation, verification, and often a formal organizational affiliation.
What is the main problem AI guardrails create for penetration testers?
The core problem is refusal on legitimate offensive security tasks. Writing proof of concept exploits, analyzing malware samples, simulating phishing content for training programs, and generating attack pattern documentation all trigger refusals in most general models. These are not theoretical use cases; they’re standard deliverables in red team engagements.
Do threat actors actually use AI for cyberattacks?
Yes, and the numbers are not small. According to SlashNext’s State of Phishing report, AI generated phishing attacks grew over 4,000% in the 18 months after large language models became widely available. Uncensored models have circulated on criminal forums since at least mid-2023. The guardrails on consumer models have not slowed this down.
Can independent security researchers get AI access without the restrictions?
Some can, but it takes work. Enterprise agreements with providers include provisions for security research use cases, but those agreements require organizational affiliation and documentation. Specialized security AI tools built for offensive research workflows are also starting to appear. The access gap is widest for solo researchers and small teams who lack the resources to negotiate directly with providers.
What does this mean for the future of AI guardrails in cybersecurity research?
Pressure from the security research community is building. Public CVE databases, bug bounty programs, and academic security research all depend on researchers being able to work with the same tools attackers use. The most likely outcome is tiered access models where verified researchers get expanded permissions, similar to how financial data providers handle institutional versus retail access.


