Instant research discovery
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Search and browse arXiv CS/AI/ML papers, enriched with AI-generated insights.
Generate novel research ideas grounded in real arXiv papers with Brainstorm.
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
Showing 331-335 of 335 papers (page 12 of 12)路 129 ms
| Paper | Published | Area | Tags | Intel | Citations |
|---|---|---|---|---|---|
| Autonomy Evaluation: Measuring LLM Agent Capabilities for Autonomous Replication and Adaptation METR (formerly ARC Evals) Published: -Area: Safety EvaluationCitations: - Tags: ai-safety, benchmark, safety-evaluation | - | Safety Evaluation | ai-safety, benchmark, safety-evaluation | E6 / R4 (96%) | - |
| BeaverTails-IT: Towards a Safety Benchmark for Evaluating Italian Large Language Models Alberto Sormani, Claudio Stamile, Daniel Scalena, Edoardo Michielon Published: -Area: Safety EvaluationCitations: 1 Tags: ai-safety, benchmark, safety-evaluation | - | Safety Evaluation | ai-safety, benchmark, safety-evaluation | E8 / R3 (97%) | 1 |
| Catch Me If You Can: Rogue AI Detection and Correction at Scale Christoph Reich Published: -Area: Deception & FailureCitations: - Tags: ai-safety, benchmark, deception-failure | - | Deception & Failure | ai-safety, benchmark, deception-failure | E4 / R3 (94%) | - |
| Evaluating MLLM Security for the Arabic and French Languages Majed Sanan, Mohamad Nassar, Razane Abdallah Published: -Area: Multimodal SafetyCitations: - Tags: adversarial-robustness, ai-safety, benchmark, multimodal-safety | - | Multimodal Safety | adversarial-robustness, ai-safety, benchmark, multimodal-safety | E2 / R1 (96%) | - |
| FlowJD: Your Imagination Can Help You Jailbreak in Visual Language Models Ke Li, Qianqian Han, Xiaotian Zou, Yongkang Chen Published: -Area: Multimodal SafetyCitations: 6 Tags: adversarial-robustness, ai-safety, benchmark, multimodal-safety, safety-evaluation | - | Multimodal Safety | adversarial-robustness, ai-safety, benchmark, multimodal-safety, safety-evaluation | E4 / R3 (95%) | 6 |