Natural Language Generation and Summarization
COMS 6975 — Fall 2026 — Columbia University
Subject to changeCourse Overview
There has been a paradigm shift in the development of models to generate language for different purposes — classification, document summarization, generation, creative writing, and image captioning to name a few. This success has largely come about due to rapid advances in large language models such as GPT5, Claude, Qwen and many others.
In this class, we will explore four main topics: language generation, multimodal generation, summarization and ethics. We will study large language models that have been used for these tasks and the issues that arise with their use. For example, how can we control the output of these large language models along different dimensions? How do we evaluate the text generated by such systems? How can we develop models to produce or summarize creative texts? What are the ethical issues surrounding these kinds of models?
We will have some invited speakers, as shown on the syllabus. Typically an invited speaker will present for half of the class and may present remotely. Starting on Sept. 30th, half the class will consist of a debate on two papers for that day. For each of the papers, one student will present the paper from a positive point of view and a second student will critique the paper.
| Time | W 4:10–6:00pm |
|---|---|
| Professor | Kathleen McKeown |
| Office Hours | W 1:00–2:00pm, Th 5:00–6:00pm (CEPSR 516) |
| kathy@cs.columbia.edu |
Teaching assistants and their office hours are listed below.
| TA | Location | Time |
|---|---|---|
| Zach Horvitz (zfh2000@columbia.edu) | Schermerhorn (room TBD) | TBD |
| Marvin Limpijankit (ml4431@columbia.edu) | Schermerhorn (room TBD) | TBD |
Requirements
Students who take the class will have four main assignments:
- For each class there will be a reading assignment consisting of several research papers. Students are responsible for reading all papers.
- Each student will be part of a debate in which they will be responsible for either presenting a paper or raising critiques about that paper. The class will participate in discussion of the paper following the debate. Class participation will be graded.
- Each student will carry out a semester-long project. This project requires submission of: a. a proposal for the project near the beginning of class; b. a midterm progress report and c. a final report and code for their project.
- Two quizzes on the class reading and debates.
Late submission policy: You have 4 free late days to use across the 3 assignments for the course (proposal, midterm, final). Once you have used all your days, then you will lose 7% of the points per day late unless you have demonstrated a valid medical excuse.
There will be no midterm or final exam.
Prerequisites
Students must have received a B or better in COMS 4705 (NLP) or equivalent. The version of NLP that you took must have covered deep learning methods for tasks in NLP.
There will be a form asking you to provide information about your background. The form also includes a short assignment required for eligibility. This form will be made available mid-August.
This is due the day after the first class (Thursday, September 10). A percentage of the class will be accepted early for forms completed before September 1. Only students who fill out the form and meet the requirements will be considered for entrance into the class — please do not request approval by email.Syllabus
This is still being finalized. Topics, speakers, and readings may change.
| Class | Date | Topic | Selected Readings |
|---|---|---|---|
| 1 | Sept 9 | Introduction to class (Kathy McKeown) LLM basics (Zach Horvitz) |
|
| 2 | Sept 16 | Summarization successes (Kathy McKeown) Evaluation Practice debate presentation by Zach Horvitz · Rebuttal by Marvin Limpijankit |
Lecture
Practice Debate
|
| 3 | Sept 23 | Frontiers of summarization: narrative (Kathy McKeown) Perspective summarization (Nick Deas) |
Lecture
|
| 4 | Sept 30 | Beyond Next Token Prediction (Zach Horvitz) Debate presentations begin |
Lecture
|
| 5 | Oct 7 | Frontiers in generation: persona vector generation / steering (Kathy McKeown) Debate |
|
| 6 | Oct 14 | Multimodality (Marvin Limpijankit) Debate |
Lecture
|
| 7 | Oct 21 | Interpretability: overview (Kathy McKeown) Debate |
Lecture
|
| 8 | Oct 28 | Industry approaches Yanda Chen (Anthropic) Debate |
Lecture
|
| 9 | Nov 4 | Interpretability: artificial impressions (Nick Deas) Debate |
Lecture
|
| 10 | Nov 11 | Ethics (Kathy McKeown) Debate |
|
| 11 | Nov 18 | Agent based approaches Elias Stengel-Eskin (UT Austin) Debate |
Lecture
|
| Nov 25 — No class (Thanksgiving) | |||
| 12 | Dec 2 | Ethics Esin Durmus (Anthropic) Debate |
Lecture
|
| 13 | Dec 9 | Industry views Kailash Karthik Saravana Kumar (Cohere) Class closing (Kathy) |
Lecture
Debate
|
Readings
There is no textbook for the class. Readings will be assigned alongside each topic on the syllabus above as the semester is finalized.
Additional readings (optional):
- Language Models Are Few-Shot Learners
- Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Self-Refine: Iterative Refinement with Self-Feedback
- Sparks of Artificial General Intelligence: Early Experiments with GPT-4
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- A Systematic Study of Position Bias in LLM-as-a-Judge
- Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
- JudgeLM: Fine-tuned Large Language Models are Scalable Judges
- SummEval: Re-evaluating Summarization Evaluation
- Large Language Diffusion Models
- No Compute Left Behind: Rethinking Reasoning and Sampling with Masked Diffusion Models
- P3Sum: Preserving Author's Perspective in News Summarization with Diffusion Language Models
- Mix and Match: Learning-free Controllable Text Generation using Energy Language Models
- COLD Decoding: Energy-based Constrained Text Generation with Langevin Dynamics
- Visual Instruction Tuning
- Representational Similarity via Interpretable Visual Concepts
- Evaluating Object Hallucination in Large Vision-Language Models
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
- Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
- Toy Models of Superposition
- Whither symbols in the era of advanced neural networks?
- LLMs model how humans induce logically structured rules
- Representation Engineering: A Top-Down Approach to AI Transparency
- LiveNewsBench: Evaluating LLM Web Search Capabilities with Freshly Curated News
- Introducing Bloom: an open source tool for automated behavioral evaluations
- Kimi K2: Open Agentic Intelligence
- SOTOPIA-S4: a user-friendly system for flexible, customizable, and large-scale social simulation
- Safety Alignment Should Be Made More Than Just a Few Tokens Deep
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- People Readily Follow Personal Advice from AI but It Does Not Improve Their Well-Being
- Responsible Evaluation of AI for Mental Health
- Extracting Training Data from Large Language Models
Academic Integrity
Copying or paraphrasing someone's work (code included), or permitting your own work to be copied or paraphrased, even if only in part, is not allowed, and will result in an automatic grade of 0 for the entire assignment or exam in which the copying or paraphrasing was done. Your grade should reflect your own work. If you believe you are going to have trouble completing an assignment, please talk to the instructor or TA in advance of the due date.