Eye Tracking 101

Before I introduce my friends at Team Gleason, let me wrap up the story of my eyegaze issues. If you recall, I struggled with accuracy on my TDG13 at first. I ended up discovering two ways to combat this issue. One was chucking the TDG13 entirely for a surface pro using a PCEye5 camera bar mounted at the bottom. The other involved a rather clever trick (if I do say so myself) involving the use of different profiles on the device. Then there was still the issue of maintaining power for the full day. 

I mentioned my experimenting with the PCEye5, but I didn't explain my theory as to why it performed better for me than the TDG13. This part gets technical, but I'll try and explain in plain terms. The short answer is that the PCEye5 has an extra IR illuminator. Duh!! Ok, ok. Let's take the helicopter tour over the volcano view of how eyegaze, or more accurately, eye tracking, works. It's actually kind of cool, but I don't want to get too bogged down with technicalities, so we'll not be diving into the volcano. Hopefully everybody has a casual understanding, or at least has heard, of the retina, pupil, and cornea as parts of the eye. If not, then I'm sorry, but your going to be looking out the wrong window and confused as to why volcanoes look like what you always thought were clouds. We don't really have to get into how the eye works, so those of you rifling through your recollections of an upside-down image on the back of the retina, can put your hands down. 

Conceptually, eye tracking is very simple. (Admittedly this would be easier with visual aids, but I'm lazy, and you'll search on it if more interest remains.) The important bit is for a camera to capture the precise location of the center of the pupil along with some fixed point, such that a line can be drawn between them representative of the line of sight. It's actually simple geometry, albeit with complicating factors to contend with. Remember that the cornea is the clear front surface of the eye, is dome shaped, and sits in front of the iris and pupil (not to be confused with the lens which is behind the pupil). The iris is the colored part of the eye and contains the muscles that dilate or constrict, changing the size of the pupil. The pupil is just a hole. It appears black because there's no light source within the eye, unless you're superman, gazerbeam, or cyclops (X-men fanatics' controversy over cyclops energy beam, notwithstanding). The retina is the light sensitive back wall, also behind the pupil. You've all seen pictures of people with "red-eye". Or maybe a dog or other animal with green or white eyes at night, as flashlights are shone on their faces. That's the light reflecting off the retina, back out. It's essentially the pupil! With special software, you can isolate that pupil and calculate its center. In eye tracking, infrared "flashlights" are used, because they are just beyond the wavelength of light that the human eye can perceive. It would be a terrible strain to have a flashlight blinding you while perusing cat videos. The wavelength is actually very close to the red end of the spectrum, so most times you will notice a distinct red glow emanating from the infrared emitters. This is referred to as near infrared (NIR). 

Let's suppose for the moment, that our camera, capable of picking up NIR light, is directly in front of our eye at some known distance. Let's further suppose that our camera is equipped with a NIR emitter surrounding it, or coming straight out from the center of our camera. We'll capture, say 60 images every second in order to keep up with a fast moving eyeball. Each image will contain that reflection off the back of the retina just like our red-eye photos from the flash. However, the NIR light will provide much better contrast between the light reflecting off the retina through the pupil than any reflecting off the iris. This provides a very clear image for computer vision software to isolate the pupil and identify its center point. There's another reflection visible too. The same NIR light will reflect off the cornea. This is referred to as a "glint". This glint from the cornea is a much smaller dot. Now, remember, that the image we're processing is a 2 dimensional representation of a 3 dimensional eyeball. We'll wave our magic maths wand of computer vision processing and draw a line between the glint spot on the cornea and the center of the pupil. This line is representative of the gaze of the eye. If the glint spot is below the pupil center, then the gaze is upwards. If the glint spot is to the left of pupil center, the gaze is to the right (assuming no horizontal mirroring is being performed, otherwise it's opposite). Our magic wand software can also create a ray, based on the simple line between our two points, that jumps out of the image into our 3 dimensional world and allows for a pretty accurate gaze direction. Where the geometry gets complicated, is working out all the angles involved as the head moves in relation to the camera, or if the NIR illuminator is offset from the camera, or if there's more than one illuminator creating another corneal glint spot and pupilary image. Calibration is necessary in order to setup parts of this geometry. Regardless, our simple scenario makes the point. So with the extra NIR illuminator, the software can compensate better and provide more accurate gaze direction and over a larger range, for larger monitor support, for example. 

I guess that was a longer helicopter ride than I thought. Since it's getting quite dense, I'll end here and pick back up and in the next one. We're getting close to some cool stuff. I promise. Thanks for hanging with me. Peace friends. 

Comments

Popular posts from this blog

He is free!

Funeral details--UPDATED

Update and thank you