Running Reproduction: MindFlow: Mind Supernet Powered Thinking Flows for Research Idea Innovation π― Explore and manage research logbooks with AI collaboration
Running Reproduction: MetaStreet: Semi-Supervised Multimodal Learning for Street-Level Socioeconomic Prediction π― Explore project logs, traces, and workspace in a web UI
Running Reproduction: Seg-ReSearch: Segmentation with Interleaved Reasoning and External Search π― Explore experiment logs and collaborate with an AI agent
Running Reproduction: IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning π― Explore research logs and collaborate with an AI agent
Running Reproduction: InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problem π―
Running Reproduction: Towards Execution-Grounded Automated AI Research π― Explore AI experiment logs and sync with your coding agent
Running Reproduction: Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models π― Browse experiment logs, traces, and workspace in a web UI
Running Reproduction: Building Social World Models with Large Language Models π― Browse experiment logs, traces, and workspace online
Running Reproduction: Auditing Sybil: Explaining Deep Lung Cancer Risk Prediction Through Generative Interventional Attributions π― Explore code, traces, and workspace with a collaborative logbook
Running Reproduction: DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation π― Explore research agent logs and traces in an interactive workspace
Running Reproduction: Judging What We Cannot Solve π― Explore code logs, traces, and workspace in a web logbook
Running Reproduction: From Knowledge to Inference: Formalizing Specialized Public Health Reasoning on GlobalHealthAtlas π― Explore code logs, traces, and workspace in an interactive logbook
Running Reproduction: SEDRAS (WZ-LLM Claims) π― Explore experiment logs, traces, and workspace in a web UI
Running Reproduction: Asymmetric Contrastive Objectives for Efficient Phenotypic Screening π― Browse logs, traces, and workspace with a collaborative web logbook
Running Reproduction: ORLoopBench: Solver-in-the-Loop Benchmarks for Self-Correction and Behavioral Rationality in Operations Research π― Explore and manage ORLoopBench benchmark logs
Running Reproduction: Seizure-Semiology-Suite (S3) π― Explore and manage seizure logbook with agent collaboration
Running Reproduction: Agentic Framework for Epidemiological Modeling π― Explore model logs and collaborate with an AI agent
Running Reproduction: An Interactive Paradigm for Deep Research π― Explore research logs, traces, and workspace interactively
Running Reproduction: Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models π― Explore and manage experiment logs with AI collaboration
Running Reproduction: CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research? π― Explore scientific experiment logs, traces, and workspace files
Running Reproduction: RL for Tool-Calling Agents in FHIR π― Explore code logs, execution traces, and workspace
Running Reproduction: Small Agent Group is the Future of Digital Health π― Explore project logs and collaborate with an AI agent