Concretely, U2ネットセグメンテーションモデルは ピクセルごとのアルファマットを予測し 背景はカットされます しかし損失源は ソフトで圧縮されたエッジを持ちます だからマットは後で 洗練されます 出力は 損失源に戻るよりも アルファチャネルを持つフォーマットに 戻らなければなりません. There is no manual masking step — a neural segmentation model does the separation and you get a transparent result back.
このサイトのコンバータはMP3、WAV、FLAC、M4A、OGG、Opusの間で音声を移動します。出力フォーマットは、始めに決めるよりも、最後まで決めておくことができるものです。 Doing that after JPG の背景を削除 rather than before means the encode happens once, on the file you actually settled on, instead of twice.