All pastes #1975516 Raw Edit

CELT discussion on #x264

public text v1 · immutable
#1975516 ·published 2010-10-28 14:32 UTC
rendered paste body
<Dark_Shikari> http://x264.nl/developers/Dark_Shikari/compare.7z
<Dark_Shikari> ABX these.
<bofh__> Hm. There's a certain distinct way the main and the backing melodies tend to influence each other in ZUN songs.
<Dark_Shikari> Rate them in order of quality, best to worst.
* _s|mon is now known as s|mon
<bofh__> rofl this is so weird. I know this is your magical overtone transform, because I listen to 1 and it sounds fine, I listen to 2 and it sounds better, it has a distinctly wider spectrum, but always whenever I listen to the original and a lowpassed second, after listening to the original the lowpassed version sounds "flat" - but 1 does not. Yet 2 still sounds better than 1. Final rating btw is 4 2 1 5 3.
<Dark_Shikari> no it isn't
<Dark_Shikari> it's 5 different encoders.
<Dark_Shikari> =p
<Dark_Shikari> no transform fuckery going on here
<Dark_Shikari> ok, so 4 + original
<bofh__> Hm. I'll guess that 1 is the original.
<Dark_Shikari> want to change your ranking now that you know I'm not filter screwing you?
<Dark_Shikari> Actually, given it's an ABX
<Dark_Shikari> I should tell you what the X is.
<Dark_Shikari> Original is 2.
<bofh__> Not greatly. 3 still sounds objectionable. 2 still sounds marginally better than 1.
<saintdev> Dark_Shikari: lemme guess, CELT?
<Dark_Shikari> saintdev: at least one might be CELT.
<saintdev> :)
<Dark_Shikari> so, now redo the test knowing 2 is original.
<bofh__> 4 sounds better than 2 consistently to me even knowing 2 is the original
<bofh__> what the shit.
<bofh__> Other than that 3 is *still* dead-last, and 1 and 5 order after 2 in no particular order.
<bofh__> They're different and have different artifacting, but I can't rate either as being worse than the other.
<Dark_Shikari> 4 sounds noisier to me than 2
<Dark_Shikari> like... a celt encode
<Dark_Shikari> So anyways, try to match them now.
<Dark_Shikari> CELT, CELT + branch with advanced prediction magic, aotuv, nero aac.
<Sean_McG> evening folks
<bofh__> Dark_Shikari: 3 and 5 are aoTuV/Nero and 1 and 4 are the CELT versions? I'd probably put 5 as aoTuV and 3 as nero, and I have no clue how the two CELT variants differ aurally or even mathematically so I can't tell obviously.
<Dark_Shikari> 3 is aotuv
<Dark_Shikari> 5 is nero
<Dark_Shikari> 1 is the old celt, 4 is celt + magic
<Sean_McG> what the blazes is aotuv?
<bofh__> Vorbis with a non-shitty psychoacoustic model.
<Sean_McG> ah.
<saintdev> Aoyumi Tuned Vorbis
<Sean_McG> sounded like a brand of dildo
<Dark_Shikari> bofh__: funny that it's the worst
<bofh__> Dark_Shikari: Vorbis sucks below 80kbit imho
<saintdev> well vorbis isn't the best at low bitrates, even with aoyumi's modifications
<Dark_Shikari> funny that celt beats it though
<bofh__> Speaking of which, what are the Nero and Vorbis bitrates/q settings here?
<Dark_Shikari> all 32
<Dark_Shikari> I chose qs to get me 32 on each one
<Dark_Shikari> 32kbps mono
<Sean_McG> yeah, Vorbis at low bitrate tends to mash low transients like drum work
<bofh__> Did Nero encode LC or HE? I still can't decide if I hear characteristic SBR-type high notes in 5 or not, I'm probably imagining them.
<Dark_Shikari> How do I tell?
<Dark_Shikari> I dunno.
<saintdev> bofh__: probably did.
<Dark_Shikari> one annoying part about this short clip is that mono-izing it really kills its quality
<bofh__> run mediainfo on the mp4 it spat out? oh if you didn't force LC at 32kbit it will use HE.
<saintdev> nero uses sbr for anything below ~64kbps
<Dark_Shikari> saintdev: it's mono though
<bofh__> yeah
<Sean_McG> wb Jarrett
<bofh__> Dark_Shikari: what's it fr--oh god is this the final stage theme from Fairy Wars?
<Dark_Shikari> lol
<checkers> Dark_Shikari: what q did you use for nero?
<Dark_Shikari> checkers: something low
<checkers> it uses HE for <0.31 and HEv2 for under 0.25(??)
<Dark_Shikari> bofh__: I picked this clip because vorbis mangled it hilariously awfully
<checkers> I'm only confident about the first number there
<Dark_Shikari> Like, it's staggering how badly it fucked up the background noise.
<Dark_Shikari> checkers: lower
<bofh__> Hey, there's a reason I consistently rated it dead-last :P
<bofh__> Also, I think I know why I prefer 4 over 1.
<bofh__> Er.
<bofh__> Over 2.
<Dark_Shikari> apparently CELT quite often beats vorbis at low bitrates.  Which is funny
<Dark_Shikari> because CELT was designed to suck at low rates
<Sean_McG> oh hey I finished Umineko ep 2 last night... Rosa's "dinner" is a lot scarier in the VN than in the anime 
<checkers> CELT is CELP for higher bitrates right?
<Dark_Shikari> no
<Dark_Shikari> SILK is CELP
<Dark_Shikari> CELT is completely novel
<Dark_Shikari> well, I mean, CELT is MDCT-based
<bofh__> So while 4 *is* noisier than 2, 2 has this really annoying high-pitched transient in the noise, somewhere around 14kHz, that gets clobbered in all the other versions. 4 just happens to artifact everything else the least.
<Dark_Shikari> that's about the only non-novel thing about it.
<checkers> silk? I thought that was the proprietary skype one?
<Dark_Shikari> checkers: not anymore proprietary
<Dark_Shikari> bofh__: interesting... the original doesn't seem to have the 14khz?!
<Dark_Shikari> oh god
<Dark_Shikari> ffmpeg must have created it
<Dark_Shikari> when downsampling
<bofh__> It's faint and very high-pitched.
<Dark_Shikari> how fucking bad do you have to be?
<bofh__> Doesn't ffmpeg's resampler use like, a good resampling algorithm?
<Dark_Shikari> dunno.
<Dark_Shikari> well, resampler maybe
<Dark_Shikari> downsampling channels, I don't know
<checkers> cool. I wonder how silk compares to celt for high bitrate voice recording (32kbit)
<bofh__> oh, how did you downsample channels? just pick the left?
<Dark_Shikari> -ac 1
<Dark_Shikari> checkers: Opus beats them both
<Dark_Shikari> aka celt/silk hybrid
<bofh__> oh. no clue if that just picks one or tries to merge.
<Sean_McG> isn't 32kbit typically overkill for just voice?
<Dark_Shikari> bofh__: merge
<bofh__> if it tries to merge then yeah that explains the obnoxious high-pitched artifact
<Dark_Shikari> bofh__: why would that happen?  I mean, you'd have to merge pretty shittily, right?
<bofh__> yes
<bofh__> What does it do to merge channels?
<bofh__> Sum both with clipping?
<checkers> that's cool, so the existing opus encoder is better than both??
<Dark_Shikari> checkers: it's just a hybrid of both encoders
<Dark_Shikari> mega hack
<bofh__> But yes. I sort of want to look into CELT now.
<Dark_Shikari> bofh__: here's the really interesting part about how CELT works
<bofh__> I'm impressed it beat both Vorbis and Nero.
<Dark_Shikari> well, a few interesting parts
* checkers listens too
<Dark_Shikari> 1) the overlap is fixed at 120 samples (WTF?!)
<Dark_Shikari> it was designed around 256 sample frame sizes, roughly, but it can extend to up to ~960 or so
<Dark_Shikari> But the overlap is fixed!
<Dark_Shikari> What's with that?  Oh, it avoids pre-echo and reduces latency.
<bofh__> Yep.
<Dark_Shikari> 2) It can subdivide each frame into smaller transforms; useful for dealing with transients + large frame size.
<Dark_Shikari> now comes the interesting parts
<Dark_Shikari> 3) It codes energy and shape separately
<Dark_Shikari> and predicts them separately!
* wewk_ is now known as wewk
<Dark_Shikari> Both are predicted in two ways
<Dark_Shikari> a) from past frames (in some fashion)
<Sean_McG> wasn't CELT or something like it used in the oldschool IP telephony apps like CU-SeeMe?
<bofh__> Wait, it does prediction?
<Dark_Shikari> no
<Dark_Shikari> to Sean_McG 
<Dark_Shikari> b) from previously-coded bands in the same frame
<Sean_McG> hm.
<Dark_Shikari> e.g. it infers higher frequencies from lower ones
<bofh__> No. You're probably thinking of a CELP/ACELP-based codec, Sean_McG.
<Dark_Shikari> Not entirely, like SBR
<Dark_Shikari> But as a prediction.
<Sean_McG> bofh__: ahhh, yeah.
<checkers> CELT is used in mumble and that's where I use it
<Dark_Shikari> energy can easily be predicted from past frames
<bofh__> CEL*P* and CEL*T* are very, very, very different things. :V
<Dark_Shikari> I mean, you couldjust guess you have the same energy as the previous frame
<Dark_Shikari> that's a pretty good guess!
<Dark_Shikari> It does this for each band.
<Dark_Shikari> It also can trade off frequency vs time resolution, separately, in each band.
<bofh__> That's actually a really clever trick, the separate coding of both energy and shape.
<Dark_Shikari> It's primarily useful for two reasons
<Dark_Shikari> 1) it retains energy (naturally)
<bofh__> Like, it's sort of obvious and at the same time not.
<Dark_Shikari> 2) prediction is improved
<Dark_Shikari> because energy is better correlated with just energy
<bofh__> i.e. you can actually *do* prediction :V
<Dark_Shikari> as opposed to trying to correlate both at the same time
<Dark_Shikari> The same probably works in video.
<Dark_Shikari> Now, more interesting things are going on.
<checkers> http://people.xiph.org/~jm/ietfcodec/ <-- is this the current 'homepage' of opus?
<Dark_Shikari> 4) As originally designed, it's hard CBR.
<Dark_Shikari> It can be adapted to not do this, but that's what it was built around.
<bofh__> Video has a lot more redundancy in time than audio does, which is why prediction already works quite well for it unlike audio.
<Dark_Shikari> The bit allocation basically works under an assumption of CBR.
<Dark_Shikari> And not just CBR overall, but constant bits per band.
<Dark_Shikari> (roughly)
<Dark_Shikari> In fact, signalling usually costs more than it's worth.
<bofh__> Oh so it's a subband codec in the end?
<Dark_Shikari> Well, every codec has bands
<Dark_Shikari> it has to define allocation to some precision
<bofh__> Okay, I see what you mean now.
<bofh__> Anyway, go on.
<Dark_Shikari> It has an interesting quantizer.
<saintdev> CELT is pretty cool :)
<Dark_Shikari> Which builds in the concept of energy maintenance
<Dark_Shikari> e.g.
<Dark_Shikari> to dequant in, say, h264
<Dark_Shikari> you input coefficients and a multiplication factor
<Dark_Shikari> and you get back dequanted coeffs
<Dark_Shikari> in celt, you input coefficients and the ENERGY OF THE BAND.
<Dark_Shikari> It dequants such that the energy matches.
<Dark_Shikari> It also has this weird thing going on with this "PVQ/CWRS" thing with a 50-dimensional entropy coder or something.
<bofh__> Man, half of this is stuff that should have been implemented in other codecs 5 years ago, and the other half is just flat-out amazing and I have to wonder how the devs came up with it b/c it's utterly brilliant.
<Dark_Shikari> I don't really get that.
<Dark_Shikari> It's probably explained in a slide somewhere.
<Dark_Shikari> Oh yeah, and there's some arithmetic coding.  The energy is arithmetic coded
<Dark_Shikari> non-adaptive.
<saintdev> isn't most of it from ideas tried out in SILK first?
<Dark_Shikari> no
<Dark_Shikari> SILK is very different
<saintdev> or was that the vorbis-replacement
<Dark_Shikari> Oh, and to do the prediction, it FDCTs past frames.
<Dark_Shikari> Ghost?
<Dark_Shikari> The interesting thing about the prediction is as follows
<Dark_Shikari> suppose you wanted to hybridize CELT with SILK
<Dark_Shikari> that is, have SILK do LFs, what it's best at
<Dark_Shikari> and CELT do HFs, what it's best at
<Dark_Shikari> you can decode the SILK part
<Dark_Shikari> and then have CELT predict from that!
<Dark_Shikari> as if the LF was actually coded by CELT.
<Dark_Shikari> Also, they used to have a pitch predictor, aka "audio motion vectors"
<Dark_Shikari> they tossed it, but apparently it's coming back or something, I'm not sure.
<Dark_Shikari> I think that was related to sample 4.
<bofh__> holy shit, prediction, actually useful in audio. Amazing.
<Dark_Shikari> Now, the really interesting thing
<saintdev> Dark_Shikari: so was it ghost that tried most of this stuff in first?
<Dark_Shikari> from all the above
<Dark_Shikari> tell me what you think CELT will suck at most.
<bofh__> I pondered the concept of pitch prediction once but deemed it too computationally costly for too little gain.
<Dark_Shikari> saintdev: dunno, I don't know shit about ghost.
<Dark_Shikari> bofh__: iirc, their original implementation was exactly that
<Sean_McG> Dark_Shikari: music?
<Dark_Shikari> I don't know its status, feel free to ask them
<Dark_Shikari> Sean_McG: no, it does great on music.
<bofh__> No, this would work great on music.
<bofh__> This would work great on music and great at low bitrates.
<Dark_Shikari> Or, at least, most music.
<Dark_Shikari> it beats the living fuck out of mp3 at ~128kbps
<Dark_Shikari> it is insanely efficient at coding HFs
<Sean_McG> that's not hard.
<Dark_Shikari> Sean_McG: it's hard with 5ms latency
<Sean_McG> it's well known that MP3 is an awful codec
<bofh__> I'm not sure if the usual suspects (snare drum, crash cymbal, ride cymbal, hi-hat) would cause it to struggle.
<Dark_Shikari> Transients are generally hard, due to its hard CBR nature
<Dark_Shikari> but:
<Dark_Shikari> a) --vbr gives it one frame of buffer 
<Dark_Shikari> which is usually enough to absorb them
<Dark_Shikari> b) it has a very very small transform size
<Sean_McG> bofh__: yeah they tend to get "muddy" easily.
<Dark_Shikari> so it doesn't "fail even at high bitrates" like mp3 does
<Dark_Shikari> it will fail on those at low bitrates
<Dark_Shikari> but so does everything else
<saintdev> bofh__: also it has horrible transient detection
<bofh__> B) is probably why it doesn't fail too much at things like castanets.wav
<Dark_Shikari> I haven't seen anything that it isn't transparent at 128kbps with.
<Dark_Shikari> But my ears suck.
<bofh__> saintdev: what, mp3? of course :V
<Dark_Shikari> The interesting thing is that which it sucks on more than you would expect.
<saintdev> but even with it's horrible transient detection it still performs well!
<saintdev> bofh__: no, CELT
<Dark_Shikari> no, bofh__, celt
<Dark_Shikari> So, the thing it fails hardest on
<Dark_Shikari> is highly tonal content
<bofh__> Hm. Distorted guitar and stuff with a broad harmonic content?
<bofh__> Oh. rofl.
<Dark_Shikari> Like, say, a chiptune mixed in with a guitar.
<Dark_Shikari> Because highly tonal content you can't cheat with energy.
<bofh__> Yes.
<Dark_Shikari> It's like the audio equivalent of anime.
<Dark_Shikari> All edges, no detail.
<bofh__> Yet anime can be encoded very well at very low bitrates with a good psy model.
<Dark_Shikari> Really?  I've never seen an anime psy model.
<Dark_Shikari> Not in a DCT codec at least!
<Sean_McG> anime doesn't tend to have quick cuts
<bofh__> I never said anime-specific, I said good :V
<Dark_Shikari> x264's psy model is nigh-useless for anime.
<Sean_McG> unless it's for the ADHD set.
<Dark_Shikari> particularly the worst case, i.e. no detail and just pure lines
<bofh__> Interesting, because x264 seems to encode anime really well even at really low bitrates. A lot of other stuff likes to choke.
<Dark_Shikari> 1) 4x4 transform
<Dark_Shikari> 2) intra prediction
<Dark_Shikari> x264 has always done well on anime, since day 1
<bofh__> I love the stuff that throws all the bits at the edges and then turns everything else into blur-o-vision.
<Sean_McG> lol
<Dark_Shikari> well, yes, that's what x264 does well on
<Dark_Shikari> but of course, if there is no "everything else"
<Dark_Shikari> there's nothing you can optimize.
<Dark_Shikari> much like a completely tonal audio file.
<Dark_Shikari> But yeah, it's funny what it fails on.
<Dark_Shikari> It fails on what you would _think_ is the easiest thing to encode.
<bofh__> Hey, I've been toying around with writing an aacPlus encoder over the past couple months. Its limiting case is something that is probably the most piss-easy thing to encode: voice.
<saintdev> bofh__: overview of CELT's attack detection
<bofh__> Also yes, but it's not tuned for tonal content.
<saintdev> Jul 31 18:55:14 <saintdev>      lol CELT's attack detection is braindead simple
<saintdev> Jul 31 19:03:23 <saintdev>      for(i = 0; i < overlap) begin[i + 1] = max( begin[i], max( abs(in[i*channels]), abs(in[i*channels + 1]));
<saintdev> Jul 31 19:03:48 <saintdev>      threshold = .4f * begin[overlap];
<saintdev> Jul 31 19:04:01 <saintdev>      that's _it_
<saintdev> Jul 31 19:04:14 <saintdev>      compare begin[] to threshold, ofc.
<saintdev> Jul 31 19:04:17 <Dark_Shikari>  lol
<saintdev> Jul 31 19:04:38 <saintdev>      like i said, braindead simple.
<bofh__> saintdev: ...rofl
<Dark_Shikari> wwwwwwwww
<Dark_Shikari> bofh__: it's obviously a placeholder for something good in the future
<Dark_Shikari> the reason being -- they want to finalize the spec sooner rather than later
<Dark_Shikari> so they're not focusing too much on encoder decision algos
<saintdev> but it still does well!
<bofh__> This thing's impressive, though.
<Dark_Shikari> Go try it.
<Dark_Shikari> Also, the code is impossible to read.
<bofh__> I'm quite amazed that it did a better job than Vorbis, too.
<Dark_Shikari> aka "jmspeex code"
<Dark_Shikari> or derf code
<Sean_McG> IOCCC candidate?
<bofh__> oh yay Jean-Marc Valin code
<Dark_Shikari> it has a few LEGENDARY lines of code that are completely impenetrable
<Dark_Shikari> and literally look like IOCCC candidates
<bofh__> him and bellard should get together and write an IOCCC entry together
<Dark_Shikari> if(l>16)val=(val>>l-16)+((val&(1<<l-16)-1)+(1<<l-16)-1>>l-16);
<bofh__> they'd win the contest forever, probably.
<Dark_Shikari> bellard has nothing on this.
<Sean_McG> OH MY GOD LOL
<Dark_Shikari>   return (_a*(_b>>shift)-(_c>>shift)+
<Dark_Shikari>    (_a*(_b&mask)+one-(_c&mask)>>shift)-1)*inv&MASK32;
<Sean_McG> I think my brain just exploded looking at that
<bofh__> That return isn't outright terrible, just a bit cryptic.
<Dark_Shikari> then you have functions that look like this:
<Dark_Shikari>     p=_u[_k+1];
<Dark_Shikari>     s=-(_i>=p);
<Dark_Shikari>     _i-=p&s;
<Dark_Shikari>     yj=_k;
<Dark_Shikari>     p=_u[_k];
<Dark_Shikari>     while(p>_i)p=_u[--_k];
<Dark_Shikari>     _i-=p;
<Dark_Shikari>     yj-=_k;
<bofh__> Also I wasn't talking about bellard's ffmpeg code, I was talking about like, tcc :P
<Dark_Shikari>     _y[j]=yj+s^s;
<Dark_Shikari> i.e. "let me write my C like perl"
<bofh__> OH GOD WHAT
<Sean_McG> HAHAHAHAHAH
<bofh__> Please tell me that had comments.
<bofh__> Please.
<Sean_McG> Dark_Shikari++
<Dark_Shikari> no.
<Dark_Shikari> it has a function comment explaining it.
<bofh__> ...
<bofh__> sdfsdgafghfdhfgdfahdgagdsf
<Dark_Shikari> But no per-line comments.
<bofh__> Yeah, IOCCC entry.
<Dark_Shikari>  /*Returns the _i'th combination of _k elements chosen from a set of size _n with associated sign bits.
<Dark_Shikari> (that wasn't the whole function, it's about twice as long as that)
<Dark_Shikari> oh, and the fixed point/floating point abstraction results in some fun stuff.
<Dark_Shikari>    return ADD16(r, MULT16_16_Q15(r, MULT16_16_Q15(y,
<Dark_Shikari>               SUB16(MULT16_16_Q15(y, 12288), 16384))));
<Dark_Shikari>    rt = ADD16(C[0], MULT16_16_Q15(n, ADD16(C[1], MULT16_16_Q15(n, ADD16(C[2],
<Dark_Shikari>               MULT16_16_Q15(n, ADD16(C[3], MULT16_16_Q15(n, (C[4])))))))));
<Dark_Shikari> return ADD16(1,MIN16(32766,ADD32(SUB16(L1,x2), MULT16_16_P15(x2, ADD32(L2, MULT16_16_P15(x2, ADD32(L3, MULT16_16_P15(L4, x2 ))))))));
<Dark_Shikari> frac = ADD16(C[0], MULT16_16_Q15(n, ADD16(C[1], MULT16_16_Q15(n, ADD16(C[2], MULT16_16_Q15(n, ADD16(C[3], MULT16_16_Q15(n, C[4]))))))));
* Sean_McG headdesks
<Dark_Shikari> return MULT16_16_P15(x, ADD32(M1, MULT16_16_P15(x, ADD32(M2, MULT16_16_P15(x, ADD32(M3, MULT16_16_P15(M4, x)))))));
<bofh__> Yeah, this reminds me of when I tried to make sense of some of speex's code
<Dark_Shikari> This is one place where C++ is justified.
<bofh__> Except literally a thousand times worse.
<saintdev> bofh__: in short, join us in #celt
<bofh__> saintdev: done and added to autojoin list :P
<checkers> that's pretty fantastic code
<checkers> the question is: is it efficient and imprenetrable, or is it like gpac
<Sean_McG> *sigh* poor gpac
<Sean_McG> take it easy folks, I'mma go watch some Lain