π© Bitcoin Red Team Update
We're continuing a large-scale security review of the Bitcoin open-source ecosystem.
25 Bitcoin developers around the world have been working on this task non-stop for 108 hours.
We scanned 501 projects and produced 7,958 findings, of which 1,280 are rated high or critical (H+C) severity.
Our goal is to provide the most useful information. Based on feedback from maintainers, we've recalibrated our severity ratings and are now stricter about what we classify as H+C.
The share of all findings rated H+C increased to 16.2%, and the share of findings with a reproducible proof of concept increased to 24.7%.
That means our accuracy is improving.
As we harvest the low-hanging fruit, we're now focusing on improving harnesses, deploying better cyber-capable models, and probing more sophisticated attacks.
Our AI spend remains consistently high.
Our cumulative spend has passed $58K as we add more inference sources. Notably, we've started using cyber models from OpenAI and Anthropic.
These models are producing incredible results, but the vast majority of our spend (74%) still goes to Kimi K3.
Kimi K3 remains the top choice among our hunters, but we've also seen a significant increase in the use of Qwen 3.8 over the past few days.
Our goal is to strengthen the Bitcoin ecosystem and increase the resilience of our infrastructure in the age of AI.
To achieve this, we're conducting the largest red-teaming campaign in Bitcoin's history.
As a Bitcoiner, I am immensely proud of this.
We're sharing all our findings with developers for free.
This would not have been possible without the generous support of organizations like @OpenSats and donors like @FPuklowski and @vik sharma.
You can help keep the Red Team alive by donating to the OpenSats Red Team Fund here: https://opensats.org/funds/red
Last but not least, our reporting rate is picking up. Thanks to our outreach team, we've reported 29.4% of our findings so far.
We're also rolling out new infrastructure to deliver these findings to projects more quickly and securely, while improving how we incorporate feedback.
Stay tuned!
We're continuing a large-scale security review of the Bitcoin open-source ecosystem.
25 Bitcoin developers around the world have been working on this task non-stop for 108 hours.
We scanned 501 projects and produced 7,958 findings, of which 1,280 are rated high or critical (H+C) severity.
Our goal is to provide the most useful information. Based on feedback from maintainers, we've recalibrated our severity ratings and are now stricter about what we classify as H+C.
The share of all findings rated H+C increased to 16.2%, and the share of findings with a reproducible proof of concept increased to 24.7%.
That means our accuracy is improving.
As we harvest the low-hanging fruit, we're now focusing on improving harnesses, deploying better cyber-capable models, and probing more sophisticated attacks.
Our AI spend remains consistently high.
Our cumulative spend has passed $58K as we add more inference sources. Notably, we've started using cyber models from OpenAI and Anthropic.
These models are producing incredible results, but the vast majority of our spend (74%) still goes to Kimi K3.
Kimi K3 remains the top choice among our hunters, but we've also seen a significant increase in the use of Qwen 3.8 over the past few days.
Our goal is to strengthen the Bitcoin ecosystem and increase the resilience of our infrastructure in the age of AI.
To achieve this, we're conducting the largest red-teaming campaign in Bitcoin's history.
As a Bitcoiner, I am immensely proud of this.
We're sharing all our findings with developers for free.
This would not have been possible without the generous support of organizations like @OpenSats and donors like @FPuklowski and @vik sharma.
You can help keep the Red Team alive by donating to the OpenSats Red Team Fund here: https://opensats.org/funds/red
Last but not least, our reporting rate is picking up. Thanks to our outreach team, we've reported 29.4% of our findings so far.
We're also rolling out new infrastructure to deliver these findings to projects more quickly and securely, while improving how we incorporate feedback.
Stay tuned!
198