We present results that shows (i) How affordance classification is enhanced by metric learning, (ii) How incremental learning affects classification, (iii) How the learned transform can ground the affordance in the feature representation and how that can be visualized on the object, and (iv) How the classification and grounding is affected by a human ordering of object similarity.
We define the affordances to be reasonable from a human perspective but also not to stretch too much from what is possible for the robot to read from its sensors. For example we use a Kinect1 camera to collect data which is quite sensitive and noisy both in 2D and depth, this limits the resolution and measurements of the point cloud of the object. We choose the following affordances
\small
- Hangable - All objects that affords hanging such as cups with ears.
- Putable - All objects that affords containing something.
- Rollable - All objects that affords rolling on the table such as cups without ears, bottles, etc.
- Lift Top - All objects that has a top that affords removing such as bottles, pens, etc.
- Tool - All objects that affords some kind of tool use this might be a spatula, hammer, screwdriver, etc.
- Brushing - All objects that affords brushing.
- Hammering - All objects that affords hammering this includes a hammer, a screwdriver, a brush, that is, objects that be replaced if a hammer is not available.
- Stirring - All objects that affords stirring again spatulas, or other objects that possibly could be used for stirring.
- Spraying - All objects that affords spraying, that is, spray bottles.
- Scraping - All objects that affords some kind of scraping action such as window scraper, etc.
- Writeable - All objects that affords writing, that is, pens.
- Drinking - All objects that affords drinking from such as bottles, cups, etc. Here we define some objects as not drinkable from social conventions or direct harm, even though they technically affords it.
- Opening - All objects that affords opening, that is all food item, boxes, bottles, etc.
- Squeezing - All objects that affords squeezing something out them such as shampoo bottles, etc.
- Playing - All objects that affords some kind of playing such as maracas, some balls and toys.
\normalsize
Our dataset consists of a set of 103 diverse instances, some are the same object but from different viewpoints, such as different tools, cups, bottles, balls, boxes, pots, cans, cleaning fluids, etc. Due to space limitations a link is provided to all images and point-clouds of the dataset.
We perform the collection of object data using a Kinect camera with objects placed on a flat surface in front of the robot. The robot observes and segments out the object and then records features over it. Each object is associated with a binary label indicating if it affords an action or not. In addition the robot might be provided with a human k-NN ordering for each of the objects in the demonstrated set, where
We preprocess the features by centering and scaling to unit variance. To evaluate the learning we run the optimization a 100 times using random splits of the collected data with a ratio of
A less expensive approach is to keep track of the ratio between the two error terms and the k-NN leave-one-out error on the training dataset. A low ratio and zero leave-one-out error almost always indicates overfitting. As for the
Affordance classification accuracy and standard deviation can be found in table \ref{fig:affordance_table1}. As we can see the accuracy is above
LMCA can use any distance measure for computing the
As mentioned in the introduction experiments by \cite{POSNER:1967ef} and others using distorted patterns have shown that the categorization process for humans happens in a continuum. In the beginning individual exemplars are remembered but as more examples are introduced a generalization process takes place. We are therefore interested in how the affordance transforms changes as more examples are introduced.
\input{fig_affordance_classification}