Posts

Showing posts from 2020

OpenAI, Week 7-8 // Timeline of AI Milestones

To round out my two month curriculum period, I thought I would share a timeline I've been assembling (with help from a fellow Scholar - thanks, Shola!) of the major deep learning research milestones from 2000 to 2020. This was a necessary step for me, as I came in having zero perspective on the field and knowing almost none of the jargon - I didn't know what ResNet, BERT, or even Transformer meant! Writing the timeline not only helped me fill in the gaps in my vocabulary, it also gave me a startling and visceral understanding of just how new all this stuff is. There's at least one major paradigm shift, breakthrough, or algorithmic revolution every year, it seems! I'm incredibly excited at the idea of contributing to this ever-growing body of research, and can't wait to start my project. 2003 - Term "word embeddings" coined 2006 - Deep Belief Networks introduced; contrastive loss introduced 2008 - GPU revolution begins 2009 - ImageNet goes live 2011 - ...

OpenAI, Week 5-6 // Implementing Transformer in PyTorch

These past two weeks, I've been studying Transformers and coding up my own implementation in PyTorch. There are quite a few excellent step-by-step guides to understanding the Transformer architecture, with gorgeous visualizations, so I'm not going to try to compete! Instead, I want to achieve two things with this post: 1) share my curated curriculum recommendations, and 2) describe a few subtleties that were not addressed in any of the tutorials I came across. 1) Curriculum As I said, there are a lot of options when it comes to learning material - a potentially overwhelming number of them. Here are my favorites, in the order I recommend accessing them: The Illustrated Transformer   by Jay Alammar Attention Is All You Need: The Transformer   by Lennart Van der Goten Walkthrough: The Transformer Architecture by Matthew Barnett Positional Encoding   by Amirhossein Kazemnejad The Annotated Transformer   by Alexander Rush 2) Subtleties Here are a couple of things I ...

OpenAI, Week 3-4 // Implementing ResNet in PyTorch

Image
Last week, I studied CNNs (mostly through Andrej Karpathy's CS231n lectures and notes). This week, I implemented ResNet end-to-end in PyTorch and trained it on the CIFAR-10 dataset. I was initially a bit confused, because the ResNet paper describes two slightly different sets of architectures: the main one optimized for ImageNet, and a narrower one optimized for CIFAR-10. I ended up writing my code such that it's flexible enough to implement both versions, depending on how it's called.  While I found many helpful guides and discussions on ResNet online, none quite laid out the details of these architectures in a "cheat sheet" way. In case it's useful to anyone else, I'm posting my handwritten notes. Below is a summary of the architecture of ResNet34. I use the notation where $K$ is the number of filters, $F$ is the size of the filter, $S$ is the stride, and $P$ is the padding.   And here is an explicit work-through of the dimensionality of this problem:...