This graduate-level course examines computational methods for understanding and generating visual data. The first part of the course develops the foundations of computer vision: image formation, light and color, sampling and filtering, local features and correspondence, image alignment, 3D reconstruction, optical flow, and tracking. The second part focuses on learned visual representations and modern vision systems, including convolutional networks, vision transformers, self-supervised and vision-language models, object detection, segmentation, video understanding, diffusion models, neural 3D representations, among others. Throughout the course, we will connect classical approaches based on physics, geometry, and optimization with contemporary learning-based methods, emphasizing their assumptions, strengths, failure modes, and evaluation. See the schedule for details.

Coursework will include homeworks, examinations, and a final project. The homeworks will provide hands-on experience with representative vision tasks and will use guided self-assessment. The midterm examinations will assess individual understanding of the classical and modern portions of the course. The final project will investigate a vision topic or application in greater depth and culminate in a written report and presentation.

The course assumes a strong working knowledge of linear algebra, probability, and Python programming. Prior experience with machine learning, signal processing, or image processing is recommended but not required. See the course logistics page for additional details.


  • Time: Tuesday/Thursday 2:30PM – 3:45PM
  • Location: Computer Science Building, W150
  • Discussion: Piazza
  • Homework: Canvas, Gradescope
  • Contact: Please post course-related questions on Piazza, where we will also post announcements. For external enquiries, personal matters, or emergencies, email the instructor.