Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
๐๏ธ
Building on HF
126.7
TFLOPS
Boning Cui
Bc-AI
155
5
70
Follow
Harley-ml's profile picture
jgfiiyc's profile picture
egepakten's profile picture
67 followers
ยท
76 following
AI & ML interests
He/Him. I like LLM's and VLM's. I work with my other friends to make stuff. We are in year 7 and we are enthusiastic about AI. We are based in Australia ๐ฆ๐บ
Recent Activity
reacted
to
mihailgribov
's
post
with ๐ฅ
about 15 hours ago
Will your AI agent tell you it was attacked? We took the same agent from our earlier experiment and added one thing: a twentieth tool, `escalate_security_incident`. The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models. Alarm rates ranged from 49% to zero. The unexpected result came from the newest model in the test, `gpt-6-astra`. Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished. That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability. A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it. Full experiment and results: https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked Quadrat-IPI dataset: https://huggingface.co/datasets/mihailgribov/quadrat-ipi Run your own model: https://github.com/mihail-gribov/quadrat-ipi-model-eval #prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents
reacted
to
mihailgribov
's
post
with ๐
about 15 hours ago
Will your AI agent tell you it was attacked? We took the same agent from our earlier experiment and added one thing: a twentieth tool, `escalate_security_incident`. The system prompt said nothing about attacks or when to use it. We then ran the same 395 injected emails through nine agentic models. Alarm rates ranged from 49% to zero. The unexpected result came from the newest model in the test, `gpt-6-astra`. Astra did not follow a single injected payment instruction. But it did not report a single one either. On clean and injected emails alike, it simply read the email, logged the subject, and finished. That is a useful distinction: resisting an attack and recognizing it as a security event are not the same capability. A model can be perfectly resistant in this test and still leave you with no evidence that anyone attacked it. Full experiment and results: https://huggingface.co/blog/mihailgribov/will-the-agent-tell-you-it-was-attacked Quadrat-IPI dataset: https://huggingface.co/datasets/mihailgribov/quadrat-ipi Run your own model: https://github.com/mihail-gribov/quadrat-ipi-model-eval #prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents
View all activity
Organizations
Bc-AI
's models
14
Sort:ย Recently updated
Bc-AI/T1-Mini-Preview-Q2_K-GGUF
Image-Text-to-Text
โข
4B
โข
Updated
8 days ago
โข
148
Bc-AI/T1-Mini-Preview
Image-Text-to-Text
โข
4B
โข
Updated
9 days ago
โข
43
โข
1
Bc-AI/Code-thing
Updated
10 days ago
Bc-AI/joint-flow-3b
Updated
12 days ago
โข
1
Bc-AI/CodVa-1-session-001
Updated
Aug 22
Bc-AI/nova-1-standard-checkpoints
Updated
Aug 11
Bc-AI/vacinova-1-checkpoints
Updated
Aug 8
Bc-AI/nova-1-large-checkpoints
Updated
Aug 2
Bc-AI/nova-1-phase2-checkpoints
Updated
Aug 2
Bc-AI/suprava-1
Updated
Jul 31
Bc-AI/codva-10b
Updated
Jul 31
โข
1
Bc-AI/codva-exp-7b
Updated
Jul 28
Bc-AI/nova-chat-ui
Updated
Jul 1
Bc-AI/sam-x-chat-instruct
Updated
Oct 19, 2025
โข
1