Project Info
Each project will be in a group of 2 students.
A core part of CS 224V is a project you work on throughout the quarter. We have mentors and advisors across many disciplines who have signed on to help you with these projects, from Stanford and from external partners. All projects are mentored and supervised on a weekly basis.
Projects from CS 224V have produced key research papers that inspired commercial deep research products (OpenAI Deep Research, Google Gemini Deep Research, Databricks Genie Deep Research Mode). The open-source software has been downloaded and used in industry, and pilots based on the technology, hosted at WWKnowledge.org, have been used by over 800K consumers, journalists, and historians. Many projects build on previous years' technology to advance the state of the art for the year after. Will your project show up on that list next year?
The project proposals document is the starting point for choosing a topic. It collects state-of-the-art LLM research projects led by NLP researchers at Stanford, together with applications of LLMs developed with experts in biomedicine, history, journalism, medicine, and sustainability. You are most welcome to propose your own project — please post it in that document so it can attract partners.
Infrastructure for Your Project
This year we are making available AOS (Agentic OS), a new architecture on which you can build your own agents. You are welcome to use any tools you like. However, your project must use state-of-the-art tools and concepts, and cannot be built with vanilla LLMs by simply prompting them with the problem statement.
AOS is an open, accountable platform for collaborative agent systems, with a layered architecture: a semantic file system for persistent knowledge and artifacts, application platforms for long-horizon interaction and accountable decision making, and team workflow support for shared workspaces and collaboration. As an open-source, model-agnostic platform, it makes smaller and open-weight models more effective, and supports data sovereignty by letting organizations deploy, audit, extend, and govern it in their own environments.
Semantic file system
Ingestion
- CHURRO: an open-weight vision-language model for historical prints and manuscripts. Trained on the largest annotated historical dataset to date, spanning 46 language clusters and 22 centuries, it is more accurate than the best commercial model at 0.6% of the cost.
Semantic retrieval
- SUQL (Structured and Unstructured Query Language): a query language that composes information from databases and textual corpora.
- SLIDERS: handles document sets of arbitrary size by structuring long documents chunk by chunk, where long-context LLMs are expensive, inaccurate, and never long enough.
- SatIR: fast, high-precision, high-recall retrieval that finds documents whose constraints are satisfied by the data in the input query (e.g. clinical trials).
Knowledge curation
- WikiChat: question answering over a text corpus.
- Co-STORM: an interactive deep researcher that writes a one-pager.
- DataTalk: an interactive question-answering interface to databases.
- DataSTORM: insight discovery across documents and databases.
- GRILL: an interactive deep researcher that can curate hundreds of papers.
Application frameworks
- Genie Worksheets: creation of long-horizon, long-context conversational task agents with a significantly higher completion rate than function calling, integrating retrieval from free text and structured data with programmable conversational agent policies.
- VERDICT: accountable decision making that returns a verifiable, faithful, consistent, and actionable rationale using formal methods. VERDICT uses LLMs to translate documents into formal SMT formulae so that theorem provers can perform the complex logical tasks. (Currently developed for clinical trial matching.)
Project Axes
There are two major approaches to defining a project: you can start with an application area, or with a technique. For the latter, you will still need a domain to test your technique on.
A. Domain driven
- What is your domain of interest?
- What is the challenge in your project? For example: human-computer interaction; precision, recall, accountability; long context; long horizon; or unknown — you may need to build the first prototype to discover the challenge.
- Which techniques are you planning to use or to improve?
- How are you planning to evaluate your solution?
Take medicine as an example. OpenEvidence's accuracy on diagnosis and treatment in the MedXpertQA benchmark is limited to 44%. How can we improve it? Simply applying RAG to the PubMed literature does not work, because each paper reports the findings of one experiment; doctors rely on guidelines, which experts create by curating many papers on a topic. That opens up a range of problems:
- Scientific biomedical discovery with experimentation (curation of knowledge). Which genes are responsible for drug resistance? Starting from wet-lab results that identify a long list of candidate genes, scour the literature to prioritize genes for further experimentation.
- Data-driven discoveries to inform new guidelines (data-driven discovery). From large numbers of patient records, can we find causal relationships between treatments and outcomes? Can we identify potential confounders from doctors' notes?
- Guidelines writing (curation of knowledge). How do we synthesize the findings across many medical publications?
- Application of guidelines (decision making). Diagnosis of rare diseases, where doctors lack experience and must apply known guidelines to patient records.
- Guidelines conformance (decision making). Can we apply guidelines to determine a treatment, or to check that a doctor's treatment conforms to them?
- Answering doctors' queries on EHRs (question answering). Why was a certain treatment or test ordered? What was the sequence of treatments? How long did diagnosis take? Were there misdiagnoses? What drug side effects were experienced?
B. Technique driven
You are welcome to improve the components of AOS or add new functionality. Some problems we have identified:
- CHURRO: no VLM today can handle Chinese newspapers typeset in the 1800s and early 1900s. Automatically synthesizing training data that teaches VLMs to handle such papers is a particularly attractive approach.
- SLIDERS: how do we support sophisticated queries over patient records?
- GRILL: how do we improve management of hypotheses and report generation? How do we perform deep analysis across domains, e.g. compiling information on rare diseases or compiling medical guidelines?
- DataSTORM and GRILL: how do we integrate the two so that literature search can be combined with data analysis?
- Genie Worksheets: can we derive worksheets from English descriptions, or from the accessibility information on websites?
- VERDICT: explore accountability in other domains — rare disease diagnosis and guideline adherence in medicine, governance compliance in finance.
Proposals and Deliverables
Mentor-written proposals
If you take a mentor-written proposal, you will submit a full project proposal on Gradescope. This is an extended version of the mentor-written proposal. It can be largely the same as what your mentor provided, but you should edit it if you are narrowing or expanding the scope and customizing it to your interests. It should have more detail on when each phase of the project will be completed, the datasets you will use, a proposed weekly schedule, and what each partner will work on.
In addition to the fields already in the mentor-written proposal, the full project proposal asks you to fill out:
Prior work:
Expected demo at the end of the quarter:
Weekly schedule:
Custom proposals
If you are doing a custom project, you will have to sign up to present it in class. Signing up is mandatory. If you sign up early, you will get more feedback, which you can use to update your final proposal.
Please still review the mentor-written proposals for examples of the level of detail we are looking for as you propose your custom project.
The full project proposal asks you to fill out:
Title:
Team Member(s):
Key Question:
Motivation:
Project description:
Mentor (if known):
Prior work:
Expected demo at the end of the quarter:
Weekly schedule:
Please use the following format: Project Template
Final Project Presentation, Poster, and Paper
At the end of the quarter we host a presentation and poster session for all final projects. Each group makes a 60-second presentation at the beginning, in our usual lecture location, and we move on to the poster session afterward. We adopt the same poster session guidelines as CS 224N. Dates are on the Schedule.
A final paper about your project is due in finals week. You should also submit your code, along with a README explaining how to run your program for a demo, at the same time as the paper. It is recommended to include a short video demo of your project along with your code.
Here is a suggested outline of the final paper:
(1) Abstract: 2-3 paragraphs summarizing the paper, including the results
(2) Introduction, which includes the motivation, main idea, and overall contribution
(3) Related work
(4) The core ideas
(5) Experimental results
(7) What you learned and future work
(8) Conclusion
(9) Appendix: Examples, Prompts that you engineered ...
We recommend using ACL style for your paper:
ACL style files
Publications from Past CS 224V Projects
2023
2024
- SUQL: Conversational Search over Structured and Unstructured Data with Large Language Models. Shicheng Liu, Jialiang Xu, Wesley Tjangnaka, Sina J. Semnani, Chen Jie Yu, Monica S. Lam. Findings of NAACL 2024.
- Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models. Yijia Shao, Yucheng Jiang, Theodore A. Kanell, Peter Xu, Omar Khattab, Monica S. Lam. NAACL 2024.
- SPAGHETTI: Open-Domain Question Answering from Heterogeneous Data Sources with Retrieval and Semantic Parsing. Heidi C. Zhang, Sina J. Semnani, Farhad Ghassemi, Jialiang Xu, Shicheng Liu, Monica S. Lam. Findings of ACL 2024.
- SPINACH: SPARQL-Based Information Navigation for Challenging Real-World Questions. Shicheng Liu*, Sina J. Semnani*, Harold Triedman, Jialiang Xu, Isaac Dan Zhao, Monica S. Lam. Findings of EMNLP 2024.
- Into the Unknown Unknowns: Engaged Human Learning through Participation in Language Model Agent Conversations. Yucheng Jiang, Yijia Shao, Dekun Ma, Sina J. Semnani, Monica S. Lam. EMNLP 2024.
2025
- Using Artificial Intelligence to Improve Empathetic Statements in Autistic Adolescents and Adults: A Randomized Clinical Trial. Koegel LK, Ponder E, Bruzzese T, Wang M, Semnani SJ, Chi N, Koegel BL, Lin TY, Swarnakar A, Lam MS. J Autism Dev Disord, 2025.
- Controllable and Reliable Knowledge-Intensive and Task-Oriented Conversational Agents with Declarative GenieWorksheets. Harshit Joshi, Shicheng Liu, James Chen, Robert Weigle, Monica S. Lam. ACL 2025.
- LEMONADE: A Large Multilingual Expert-Annotated Abstractive Event Dataset for the Real World. Sina J. Semnani, Pingyue Zhang, Wanyue Zhai, Haozhuo Li, Ryan Beauchamp, Trey Billing, Katayoun Kishi, Manling Li, Monica S. Lam. Findings of ACL 2025.
- CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition. Sina J. Semnani, Han Zhang, Xinyan He, Merve Tekgürler, Monica S. Lam. EMNLP 2025.
- Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models. Sina J. Semnani, Jirayu Burapacheep, Arpandeep Khatua, Thanawan Atchariyachanvanit, Zheng Wang, Monica S. Lam. EMNLP 2025.
2026
- SatIR: Scalable High-Recall Constraint-Satisfaction-Based Information Retrieval for Clinical Trials Matching. Zikai Zhou, Yufei Jin, Yilin Xu, Yu-Chiang Wang, Chieh-Ju Chao, Monica S. Lam. COLM 2026.
- DataSTORM: Deep Research on Large-Scale Databases using Exploratory Data Analysis and Data Storytelling. Shicheng Liu, Yucheng Jiang, Sajid Farook, Camila Nicollier Sanchez, David Fernando Castro Pena, Monica S. Lam. COLM 2026.
- Accountable AI with Grounded, Faithful, Consistent, Actionable Rationale: A Case Study in Clinical Trial Matching with VERDICT. Zikai Zhou, Yufei Jin, Yilin Xu, Yu-Chiang Wang, Chieh-Ju Chao, Monica S. Lam. EMNLP 2026.