I really like the bibliography in that document.
No, no, no, you haven’t understand what this paper say !!!
It only say that the Nvidia hardware have a very bad vertex cache for more than 50 vertices …
And it say too that with less than 50 vertices per instance, it can bee more that 35x more speed on some cards 
This isn’t the same thing that to say that instancing cannot be a good thing for OpenGL …