Memory and Speed issues

To put the binary vs. ASCII discussion to rest:

I posted a test scene with 1 million triangles for performance testing on the private forum (this forum doesn’t seem to accept attachments).

Interesting stats:

  • original binary file size: 74.33 MB (22.41 MB compressed)
  • COLLADA file size: 53.22 MB (9.67MB compressed) This would be even smaller without the 2-space indenting…

The COLLADA file is smaller, because small index numbers and even floats only take up 2-3 bytes, instead of the fixed 4 (or 8 for doubles) bytes in a binary file.

For a speed reference:
EQUINOX-3D loads the COLLADA file in 4-5 seconds on a 3GHz P4 Linux box (~3 sec for the binary file).
This includes processing the data for the EQX internals, such as:

  • generating a unique edge-list for wireframe rendering (1.5M edges)
  • computing triangle normals
    etc.

The actual COLLADA parsing is <3 seconds.

If anyone has a chance to convert it to a Maya/XSI/MAX binary file and give us size and import/export speed numbers I’d love to hear them.

Gabor

I gave up on the binary thing but since you posted I’m going to have to say it doesn’t seem like your data is valid.

I’ve never seen ascii floats be 2 or 3 bytes in any real data. ASCII Floats are generally at least 7-8 bytes. 1.2345 plus the separator plus the minus sign. Not that also assumes values -9.9999 to 9.9999. As soon as you go over 10 it becomes 8-9 bytes, over 100 it becomes 9-10 bytes. If most of your data your floats only took 3 bytes (ie, 1.2) then it suggests your data is not real world data.

I also know of no game based formats that use doubles for storage. In fact I don’t think any DCC’s tools I know of use doubles so the “8 bytes for doubles” arguement is also not real world and, if you were going to use doubles then you’d need to make your ascii representations longer (more percision) otherwise why start with doubles in the first place.

As for times for Maya/XSI/Max, those are irrelavent to the argument. That issue is what happens after we get the data out from one of those tools and start processing into game data. We can calculate that 1 million verts with X,Y,Z,NX,NY,NZ,U,V,R,G,B,A per vertex that’s 48meg. If those same floats were plain ascii with typical 4 point percision they would take more than 90meg (7.5 bytes per float per vert). Just in terms of loading, saving and or transfering we know that its going to take nearly twice as long even before parsing. On top of which, if you load the whole file at once and then parse into floats you need 126meg of memory. If you load the floats directly you only need 48. If you parse as you load then your parsing will slow down.

Now, multiply those times and sizes by thousands of assets and you can see where binary might have its place.

But, as for why I dropped the argument, it’s because I see the point that processing binary in XML is not easy to do outside of C++. Many other languages are not so good at loading and manipulating binary data period. (although I will point out the spec says it assumes C++).

You seemed to miss the point that it’s the integer indices that account for the majority of the size savings, not the floats.

If you have meshes with large vertex arrays, they will require 3-4 bytes for all indices in an uncompressed binary file, regardless the value of the index.
On the other hand, ASCII acts a lot like Huffmann encoding or LZW:
The smaller the value (number of significant bits), the less space it takes. E.g.:

1 2 3 4</p>
This is <2 bytes per index (7 bytes for 4 indices) and it only goes up as needed to represent the increasing numbers.

Next gen content will most likely have large meshes and you want to avoid splitting them up, because that introduces contex-switching over-head.

Cheers,

Gabor

If you check your example file you’ll see that 1 million triangles would most likely have indices from 0 to at least 500000. Any indice over 999 would take 5 bytes per. Example 1000<space>. That’s 20% more space per int and since the range on indices 0-999 is only 99.8% of the indices in a range of 0-500000 that means 99.8% of all indices. After 99999 it takes 6 bytes per or 50% more space. Assuming no padding except separators storing just an array of ints from 0-500000 in ascii would take 3388890 bytes, in binary only 2000000 bytes. That’s 69% more space for the ascii version.

Hello,

Thanks for the feedback, but I was not trying to prove that ASCII files are always smaller than binary (that would be nonsense), but sometimes it can be.
:wink:
I was pretty surprised by the results myself!
But no matter what I tried, quite often, the COLLADA file ends up being smaller.

Maybe it’s the kind of content I tried and maybe there are more efficient binary formats out there than the one I’m using, although it’s pretty compact as far as I know.
That’s why I asked to see the Maya/XSI/Max binary version of the file and other examples.
It’s nice to discuss how things should and should not be, but more actual test results would be a bit more useful.
:wink:

Cheers,

Gabor

Actually, Maya does internally store things as doubles (in both memory and the binary file format). Of course this is an explanation for why Maya needs lots of memory to run well. Collada doesn’t store the full precision that Maya uses internally itself–ASCII floats have 6 places after the decimal whereas Maya stores 8 decimal places in its “ASCII” format (counting places before and after the decimal point, trailing zeroes removed). It seems there may be some inconsistencies if you do a Maya to Collada conversion and back, or even a conversion between Maya binary and ASCII (although probably not noticeable). In the GUI, Maya usually only displays 3 places after the decimal (although I imagine this is configurable).

I did some tests of the size of a file (converted an existing game asset to Collada, then read it back in and stored it as a Maya ASCII and Maya Binary file):

Collada = 1,195,713 bytes
Maya bin = 1,009,500 bytes
Maya asc = 1,845,599 bytes

As you can see, the collada file isn’t much bigger than Maya’s binary format, and quite a bit smaller than Maya’s ASCII format (and Collada could be even smaller). The average size of a Collada float in my test is 11 bytes for positions, 10.3 bytes for normals, 10 bytes for UVs, 7.1 bytes for colors (lots of “alpha=1” values in color). The average size of a Collada index list, counting polygon markers and whitespace, was 4.9 bytes/index (although each mesh only has indices up to around 1000, with many meshes).

Just wanted to throw out some numbers…

As for my comment that Collada could be smaller – it seems it would be possible to optionally use hexadecimal values for integers rather than decimal (even floats could potentially be stored using some sort of hexadecimal encoding of mantissa / exponent). Although it would make the files much less readable I suppose. If you didn’t care about readability of number arrays at all, you could probably devise a format that used a large swath of the alphabet to encode numbers (say base 64) to generally compress the format, likely without increasing decoding time.

Jason Hoerner
Technical Director
Heavy Iron Studios, THQ Inc.

Thank you for the examples.

Gabor Nagy
www.equinox3d.com
Founder/developer