From mboxrd@z Thu Jan  1 00:00:00 1970
Return-Path: <ffmpeg-devel-bounces@ffmpeg.org>
Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org [79.124.17.100])
	by master.gitmailbox.com (Postfix) with ESMTPS id 12C714D76D
	for <ffmpegdev@gitmailbox.com>; Sat, 19 Apr 2025 20:35:11 +0000 (UTC)
Received: from [127.0.1.1] (localhost [127.0.0.1])
	by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id ADE46687DB6;
	Sat, 19 Apr 2025 23:35:05 +0300 (EEST)
Received: from mail-wr1-f42.google.com (mail-wr1-f42.google.com
 [209.85.221.42])
 by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 326D6687CEF
 for <ffmpeg-devel@ffmpeg.org>; Sat, 19 Apr 2025 23:34:59 +0300 (EEST)
Received: by mail-wr1-f42.google.com with SMTP id
 ffacd0b85a97d-39c266c2dd5so2922945f8f.3
 for <ffmpeg-devel@ffmpeg.org>; Sat, 19 Apr 2025 13:34:59 -0700 (PDT)
DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
 d=jkqxz-net.20230601.gappssmtp.com; s=20230601; t=1745094898; x=1745699698;
 darn=ffmpeg.org; 
 h=content-transfer-encoding:in-reply-to:content-language:references
 :to:from:subject:user-agent:mime-version:date:message-id:from:to:cc
 :subject:date:message-id:reply-to;
 bh=EfWMOpmk7IhlU+bLluDQIRXqtTq3zu8qReQtDG4uTf4=;
 b=js97JU0Sx/A2IwjFkT6PVqaBv0toqmiSekeTLYn0nXfLKDkZq5rd84W+ayf87ZUeg0
 Z+Iru80FzsfQ4eh/FyMyyLjtDaJb6fFe8inmkZ+8ic6mKNhe6BKoKajqcBND0lmOIKeI
 eDvcxTHM85r73yuGjFUlnawuo379BkaJ3Ru6tcjuKL/C/W+iZehVw7HFTmVWIupKCTVt
 jgmUMfCfO9L26TikWeblAXOvi64cc4S339lw+BcUGMv3T0PzgRyNZupYfmeSdjJOnxti
 sAyp+a+DHioQWC+ongt0qelKEqRApgWy7PY8DMJ1GTHPV7iJEwxdUpm0hM8+aLEDShOJ
 TlNw==
X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;
 d=1e100.net; s=20230601; t=1745094898; x=1745699698;
 h=content-transfer-encoding:in-reply-to:content-language:references
 :to:from:subject:user-agent:mime-version:date:message-id
 :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to;
 bh=EfWMOpmk7IhlU+bLluDQIRXqtTq3zu8qReQtDG4uTf4=;
 b=nbIR4OaiTeVWppEe/TjEJdeXuNCr+Bq6axaB1Q8XXdkpsaW4oDXVx3Y7r32TWK1aN2
 +h+n0yH91uUe6SYvN/RmyvjeI3HZGjm6rHIo0FwUNpYqz4l7Rj5c4sVCZta6IAkkPYa7
 QaOKTgZiT/w7FHu3SygX3f3b0OCIWAbYH6a3Vu+Mj7FQ5VP596iXqoD9eWptRe+4qH6f
 DEr/I77b6rvv9gluMh1L0RbNRk5zsIF6RfYbXjRLSdTglgu6FYDn5YgWdPXp9gxT2x5O
 bzvdIKrdrb36euKw9/+ilp38CtPRgKFyCYmVwOV6GYRw7YQhpQm7Uot6qy0REANHugxD
 41WQ==
X-Gm-Message-State: AOJu0YyjGuSIYiVN2nI4/E1qOmbv5wtexedAD/hnqoqafDuRWJ4w2x03
 MXxZSF1M+3fOgL3xDzlaS4r02naarANCE6KQIOhjqyIXoz39A20OOtAobtebeHLWI42H4HbxsS8
 b
X-Gm-Gg: ASbGncvoZGqOB5aVBePgT2ODdPeZB5FwjU6BkYdaMfJPfhRt1HMbVPS8y7LuFiDFz7x
 sVxjpgtWWlkIl7miWDgfLlQAgAqViy6B917IWPycdJIF7Cl1sco55M7NBn0/Nl/MrBLVcoUW9cU
 uuRI/u8bTCrJbVzLPYqkc5rkiICwKdjGZgx2RChxatURP6T1OWYcnuyoqUVyj1mxpjXc6fCJjps
 puAqLNKE0yzv+QoXdZWrh3cGAVWQsVJVVUyGIWtZJdhN43IfsYW6OLmKF6uN4yi2UBuQhufrpxp
 Tg3k+I/EO9Xtr0+SoF1hUzTtrUTHJcyXHchaKPowKRt+co09NGndnW3ZA2nTJQ9k87Sk0EnSEEb
 eIsIupVGawEW6EA==
X-Google-Smtp-Source: AGHT+IF4AvtCQxwATeTfG/kd9Y+NEq3Db0dgWz2jss4R7rSySDAxzV4Qsypi/cDUJ+aToJTa2O03eA==
X-Received: by 2002:a05:6000:2903:b0:39c:30d9:3b5c with SMTP id
 ffacd0b85a97d-39efbad53f0mr5442573f8f.39.1745094898190; 
 Sat, 19 Apr 2025 13:34:58 -0700 (PDT)
Received: from [192.168.0.15]
 (cpc92320-cmbg19-2-0-cust719.5-4.cable.virginm.net. [82.13.66.208])
 by smtp.gmail.com with ESMTPSA id
 ffacd0b85a97d-39efa4a4f2dsm6893698f8f.95.2025.04.19.13.34.57
 for <ffmpeg-devel@ffmpeg.org>
 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128);
 Sat, 19 Apr 2025 13:34:57 -0700 (PDT)
Message-ID: <2b794e37-9a5f-4904-bbaa-e4e999f69a48@jkqxz.net>
Date: Sat, 19 Apr 2025 21:34:59 +0100
MIME-Version: 1.0
User-Agent: Mozilla Thunderbird
From: Mark Thompson <sw@jkqxz.net>
To: ffmpeg-devel@ffmpeg.org
References: <20250419190712.1265201-1-sw@jkqxz.net>
 <20250419190712.1265201-6-sw@jkqxz.net>
Content-Language: en-US
In-Reply-To: <20250419190712.1265201-6-sw@jkqxz.net>
Subject: Re: [FFmpeg-devel] [PATCH 5/6] lavc/apv: AVX2 transquant for x86-64
X-BeenThere: ffmpeg-devel@ffmpeg.org
X-Mailman-Version: 2.1.29
Precedence: list
List-Id: FFmpeg development discussions and patches <ffmpeg-devel.ffmpeg.org>
List-Unsubscribe: <https://ffmpeg.org/mailman/options/ffmpeg-devel>,
 <mailto:ffmpeg-devel-request@ffmpeg.org?subject=unsubscribe>
List-Archive: <https://ffmpeg.org/pipermail/ffmpeg-devel>
List-Post: <mailto:ffmpeg-devel@ffmpeg.org>
List-Help: <mailto:ffmpeg-devel-request@ffmpeg.org?subject=help>
List-Subscribe: <https://ffmpeg.org/mailman/listinfo/ffmpeg-devel>,
 <mailto:ffmpeg-devel-request@ffmpeg.org?subject=subscribe>
Reply-To: FFmpeg development discussions and patches <ffmpeg-devel@ffmpeg.org>
Content-Type: text/plain; charset="us-ascii"
Content-Transfer-Encoding: 7bit
Errors-To: ffmpeg-devel-bounces@ffmpeg.org
Sender: "ffmpeg-devel" <ffmpeg-devel-bounces@ffmpeg.org>
Archived-At: <https://master.gitmailbox.com/ffmpegdev/2b794e37-9a5f-4904-bbaa-e4e999f69a48@jkqxz.net/>
List-Archive: <https://master.gitmailbox.com/ffmpegdev/>
List-Post: <mailto:ffmpegdev@gitmailbox.com>

On 19/04/2025 20:07, Mark Thompson wrote:
> Typical checkasm result on Alder Lake:
> 
> decode_transquant_8_c:                                 408.7 ( 1.00x)
> decode_transquant_8_avx2:                               94.2 ( 4.34x)
> decode_transquant_10_c:                                413.1 ( 1.00x)
> decode_transquant_10_avx2:                              87.5 ( 4.72x)

Interestingly this is signficantly better on an older Skylake.  Typical result there:

decode_transquant_8_c:                                 892.5 ( 1.00x)
decode_transquant_8_avx2:                              143.4 ( 6.22x)
decode_transquant_10_c:                                938.7 ( 1.00x)
decode_transquant_10_avx2:                             137.9 ( 6.81x)

> ---
>  libavcodec/apv_dsp.c          |   4 +
>  libavcodec/apv_dsp.h          |   2 +
>  libavcodec/x86/Makefile       |   2 +
>  libavcodec/x86/apv_dsp.asm    | 243 ++++++++++++++++++++++++++++++++++
>  libavcodec/x86/apv_dsp_init.c |  41 ++++++
>  tests/checkasm/Makefile       |   1 +
>  tests/checkasm/apv_dsp.c      | 113 ++++++++++++++++
>  tests/checkasm/checkasm.c     |   3 +
>  tests/checkasm/checkasm.h     |   1 +
>  9 files changed, 410 insertions(+)
>  create mode 100644 libavcodec/x86/apv_dsp.asm
>  create mode 100644 libavcodec/x86/apv_dsp_init.c
>  create mode 100644 tests/checkasm/apv_dsp.c
> 
> diff --git a/libavcodec/apv_dsp.c b/libavcodec/apv_dsp.c
> index 6b8b6ce590..108a3cd370 100644
> --- a/libavcodec/apv_dsp.c
> +++ b/libavcodec/apv_dsp.c
> @@ -133,4 +133,8 @@ static void apv_decode_transquant_c(void *output,
>  av_cold void ff_apv_dsp_init(APVDSPContext *dsp)
>  {
>      dsp->decode_transquant = apv_decode_transquant_c;
> +
> +#if ARCH_X86_64
> +    ff_apv_dsp_init_x86_64(dsp);
> +#endif
>  }
> diff --git a/libavcodec/apv_dsp.h b/libavcodec/apv_dsp.h
> index 13f401c4cc..a7b5377357 100644
> --- a/libavcodec/apv_dsp.h
> +++ b/libavcodec/apv_dsp.h
> @@ -34,4 +34,6 @@ typedef struct APVDSPContext {
>  
>  void ff_apv_dsp_init(APVDSPContext *dsp);
>  
> +void ff_apv_dsp_init_x86_64(APVDSPContext *dsp);
> +
>  #endif /* AVCODEC_APV_DSP_H */
> diff --git a/libavcodec/x86/Makefile b/libavcodec/x86/Makefile
> index 5d53515381..821c410a0f 100644
> --- a/libavcodec/x86/Makefile
> +++ b/libavcodec/x86/Makefile
> @@ -44,6 +44,7 @@ OBJS-$(CONFIG_ADPCM_G722_DECODER)      += x86/g722dsp_init.o
>  OBJS-$(CONFIG_ADPCM_G722_ENCODER)      += x86/g722dsp_init.o
>  OBJS-$(CONFIG_ALAC_DECODER)            += x86/alacdsp_init.o
>  OBJS-$(CONFIG_APNG_DECODER)            += x86/pngdsp_init.o
> +OBJS-$(CONFIG_APV_DECODER)             += x86/apv_dsp_init.o
>  OBJS-$(CONFIG_CAVS_DECODER)            += x86/cavsdsp.o
>  OBJS-$(CONFIG_CFHD_DECODER)            += x86/cfhddsp_init.o
>  OBJS-$(CONFIG_CFHD_ENCODER)            += x86/cfhdencdsp_init.o
> @@ -149,6 +150,7 @@ X86ASM-OBJS-$(CONFIG_ADPCM_G722_DECODER) += x86/g722dsp.o
>  X86ASM-OBJS-$(CONFIG_ADPCM_G722_ENCODER) += x86/g722dsp.o
>  X86ASM-OBJS-$(CONFIG_ALAC_DECODER)     += x86/alacdsp.o
>  X86ASM-OBJS-$(CONFIG_APNG_DECODER)     += x86/pngdsp.o
> +X86ASM-OBJS-$(CONFIG_APV_DECODER)      += x86/apv_dsp.o
>  X86ASM-OBJS-$(CONFIG_CAVS_DECODER)     += x86/cavsidct.o
>  X86ASM-OBJS-$(CONFIG_CFHD_ENCODER)     += x86/cfhdencdsp.o
>  X86ASM-OBJS-$(CONFIG_CFHD_DECODER)     += x86/cfhddsp.o
> diff --git a/libavcodec/x86/apv_dsp.asm b/libavcodec/x86/apv_dsp.asm
> new file mode 100644
> index 0000000000..0329089f45
> --- /dev/null
> +++ b/libavcodec/x86/apv_dsp.asm
> @@ -0,0 +1,243 @@
> +;************************************************************************
> +;* This file is part of FFmpeg.
> +;*
> +;* FFmpeg is free software; you can redistribute it and/or
> +;* modify it under the terms of the GNU Lesser General Public
> +;* License as published by the Free Software Foundation; either
> +;* version 2.1 of the License, or (at your option) any later version.
> +;*
> +;* FFmpeg is distributed in the hope that it will be useful,
> +;* but WITHOUT ANY WARRANTY; without even the implied warranty of
> +;* MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the GNU
> +;* Lesser General Public License for more details.
> +;*
> +;* You should have received a copy of the GNU Lesser General Public
> +;* License along with FFmpeg; if not, write to the Free Software
> +;* 51, Inc., Foundation Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA
> +;******************************************************************************
> +
> +%include "libavutil/x86/x86util.asm"
> +
> +SECTION .text
> +
> +align 32
> +const tmatrixh
> +    dw  64,  89,  84,  75,  64,  50,  35,  18
> +    dw  64,  75,  35, -18, -64, -89, -84, -50
> +    dw  64,  50, -35, -89, -64,  18,  84,  75
> +    dw  64,  18, -84, -50,  64,  75, -35, -89
> +    dw  64, -18, -84,  50,  64, -75, -35,  89
> +    dw  64, -50, -35,  89, -64, -18,  84, -75
> +    dw  64, -75,  35,  18, -64,  89, -84,  50
> +    dw  64, -89,  84, -75,  64, -50,  35, -18
> +const tmatrixv
> +    dw  64,  89,  84,  75,  64,  50,  35,  18
> +    dw  64, -18, -84,  50,  64, -75, -35,  89
> +    dw  64,  75,  35, -18, -64, -89, -84, -50
> +    dw  64, -50, -35,  89, -64, -18,  84, -75
> +    dw  64,  50, -35, -89, -64,  18,  84,  75
> +    dw  64, -75,  35,  18, -64,  89, -84,  50
> +    dw  64,  18, -84, -50,  64,  75, -35, -89
> +    dw  64, -89,  84, -75,  64, -50,  35, -18
> +
> +; Memory targets for vpbroadcastd (register version requires AVX512).
> +const one
> +    dd   1
> +const sixtyfour
> +    dd  64
> +
> +; void ff_apv_decode_transquant_avx2(void *output,
> +;                                    ptrdiff_t pitch,
> +;                                    const int16_t *input,
> +;                                    const int16_t *qmatrix,
> +;                                    int64_t bit_depth,
> +;                                    int64_t qp_shift);
> +
> +INIT_YMM avx2
> +
> +cglobal apv_decode_transquant, 6, 6, 16, output, pitch, input, qmatrix, bit_depth, qp_shift
> +
> +    ; Load input and dequantise
> +
> +    lea       rax, [bit_depthq - 2]
> +    movq      xm8, qp_shiftq
> +    movq      xm9, rax
> +    vpbroadcastd  m10, [one]
> +    vpslld    m10, m10, xm9
> +    vpsrld    m10, m10, 1
> +
> +    ; m8  = scalar qp_shift
> +    ; m9  = scalar bd_shift
> +    ; m10 = vector 1 << (bd_shift - 1)
> +    ; m11 = qmatrix load
> +%macro LOAD_AND_DEQUANT 2 ; (xmm input, constant offset)
> +    vpmovsxwd m%1, [inputq   + %2]
> +    vpmovsxwd m11, [qmatrixq + %2]
> +    vpmulld   m%1, m%1, m11
> +    vpslld    m%1, m%1, xm8
> +    vpaddd    m%1, m%1, m10
> +    vpsrad    m%1, m%1, xm9
> +    vpackssdw m%1, m%1, m%1
> +%endmacro
> +
> +    LOAD_AND_DEQUANT 0, 0x00
> +    LOAD_AND_DEQUANT 1, 0x10
> +    LOAD_AND_DEQUANT 2, 0x20
> +    LOAD_AND_DEQUANT 3, 0x30
> +    LOAD_AND_DEQUANT 4, 0x40
> +    LOAD_AND_DEQUANT 5, 0x50
> +    LOAD_AND_DEQUANT 6, 0x60
> +    LOAD_AND_DEQUANT 7, 0x70
> +
> +    ; Transform columns
> +    ; This produces the output in rows to avoid needing any
> +    ; transpose operations
> +
> +    vpunpcklwd  m12, m0, m1
> +    vpunpcklwd  m13, m2, m3
> +    vpunpcklwd  m14, m4, m5
> +    vpunpcklwd  m15, m6, m7
> +
> +%macro TRANS_COL_STEP 2
> +    vpbroadcastd   m8,  [tmatrixh + %2 + 0x00]
> +    vpbroadcastd   m9,  [tmatrixh + %2 + 0x04]
> +    vpbroadcastd   m10, [tmatrixh + %2 + 0x08]
> +    vpbroadcastd   m11, [tmatrixh + %2 + 0x0c]
> +    vpmaddwd  m8,  m8,  m12
> +    vpmaddwd  m9,  m9,  m13
> +    vpmaddwd  m10, m10, m14
> +    vpmaddwd  m11, m11, m15
> +    vpaddd    m8,  m8,  m9
> +    vpaddd    m10, m10, m11
> +    vpaddd    m%1, m8,  m10
> +%endmacro
> +
> +    TRANS_COL_STEP 0, 0x00
> +    TRANS_COL_STEP 1, 0x10
> +    TRANS_COL_STEP 2, 0x20
> +    TRANS_COL_STEP 3, 0x30
> +    TRANS_COL_STEP 4, 0x40
> +    TRANS_COL_STEP 5, 0x50
> +    TRANS_COL_STEP 6, 0x60
> +    TRANS_COL_STEP 7, 0x70
> +
> +    ; Mid-transform normalisation
> +    ; Note that outputs here are fitted to 16 bits
> +
> +    vpbroadcastd  m8, [sixtyfour]
> +
> +%macro NORMALISE 1
> +    vpaddd    m%1, m%1, m8
> +    vpsrad    m%1, m%1, 7
> +    vpackssdw m%1, m%1, m%1
> +    vpermq    m%1, m%1, q3120
> +%endmacro
> +
> +    NORMALISE 0
> +    NORMALISE 1
> +    NORMALISE 2
> +    NORMALISE 3
> +    NORMALISE 4
> +    NORMALISE 5
> +    NORMALISE 6
> +    NORMALISE 7
> +
> +    ; Vertical transform
> +
> +    mova      m12, [tmatrixv + 0x00]
> +    mova      m13, [tmatrixv + 0x20]
> +    mova      m14, [tmatrixv + 0x40]
> +    mova      m15, [tmatrixv + 0x60]
> +
> +%macro TRANS_ROW_STEP 1
> +    vperm2i128  m%1, m%1, m%1, 1
> +    vpmaddwd  m8,  m%1, m12
> +    vpmaddwd  m9,  m%1, m13
> +    vpmaddwd  m10, m%1, m14
> +    vpmaddwd  m11, m%1, m15
> +    vphaddd   m8,  m8,  m9
> +    vphaddd   m10, m10, m11
> +    vphaddd   m%1, m8,  m10
> +%endmacro
> +
> +    TRANS_ROW_STEP 0
> +    TRANS_ROW_STEP 1
> +    TRANS_ROW_STEP 2
> +    TRANS_ROW_STEP 3
> +    TRANS_ROW_STEP 4
> +    TRANS_ROW_STEP 5
> +    TRANS_ROW_STEP 6
> +    TRANS_ROW_STEP 7
> +
> +    ; Renormalise, clip and store output
> +
> +    vpbroadcastd  m14, [one]
> +    mov       rax, 20
> +    sub       rax, bit_depthq ; bd_shift
> +    movq      xm9, rax
> +    dec       rax
> +    movq      xm13, rax
> +    vpslld    m8, m14, xm13
> +    movq      xm13, bit_depthq
> +    vpslld    m12, m14, xm13
> +    vpsrld    m10, m12, 1
> +    vpsubd    m12, m12, m14
> +    vpxor     m11, m11, m11
> +
> +    cmp       bit_depthq, 8
> +    jne       store_10
> +
> +    ; m8  = vector 1 << (bd_shift - 1)
> +    ; m9  = scalar bd_shift
> +    ; m10 = vector 1 << (bit_depth - 1)
> +    ; m11 = zero
> +    ; m12 = vector (1 << bit_depth) - 1
> +
> +%macro NORMALISE_AND_STORE_8 1
> +    vpaddd    m%1, m%1, m8
> +    vpsrad    m%1, m%1, xm9
> +    vpaddd    m%1, m%1, m10
> +    vextracti128  xm13, m%1, 0
> +    vextracti128  xm14, m%1, 1
> +    vpackusdw xm%1, xm13, xm14
> +    vpackuswb xm%1, xm%1, xm%1
> +    movq      [outputq], xm%1
> +    add       outputq, pitchq
> +%endmacro
> +
> +    NORMALISE_AND_STORE_8 0
> +    NORMALISE_AND_STORE_8 1
> +    NORMALISE_AND_STORE_8 2
> +    NORMALISE_AND_STORE_8 3
> +    NORMALISE_AND_STORE_8 4
> +    NORMALISE_AND_STORE_8 5
> +    NORMALISE_AND_STORE_8 6
> +    NORMALISE_AND_STORE_8 7
> +
> +    RET
> +
> +store_10:
> +
> +%macro NORMALISE_AND_STORE_10 1
> +    vpaddd    m%1, m%1, m8
> +    vpsrad    m%1, m%1, xm9
> +    vpaddd    m%1, m%1, m10
> +    vpmaxsd   m%1, m%1, m11
> +    vpminsd   m%1, m%1, m12
> +    vextracti128  xm13, m%1, 0
> +    vextracti128  xm14, m%1, 1
> +    vpackusdw xm%1, xm13, xm14
> +    mova      [outputq], xm%1
> +    add       outputq, pitchq
> +%endmacro
> +
> +    NORMALISE_AND_STORE_10 0
> +    NORMALISE_AND_STORE_10 1
> +    NORMALISE_AND_STORE_10 2
> +    NORMALISE_AND_STORE_10 3
> +    NORMALISE_AND_STORE_10 4
> +    NORMALISE_AND_STORE_10 5
> +    NORMALISE_AND_STORE_10 6
> +    NORMALISE_AND_STORE_10 7
> +
> +    RET
> diff --git a/libavcodec/x86/apv_dsp_init.c b/libavcodec/x86/apv_dsp_init.c
> new file mode 100644
> index 0000000000..f7f209cea8
> --- /dev/null
> +++ b/libavcodec/x86/apv_dsp_init.c
> @@ -0,0 +1,41 @@
> +/*
> + * This file is part of FFmpeg.
> + *
> + * FFmpeg is free software; you can redistribute it and/or
> + * modify it under the terms of the GNU Lesser General Public
> + * License as published by the Free Software Foundation; either
> + * version 2.1 of the License, or (at your option) any later version.
> + *
> + * FFmpeg is distributed in the hope that it will be useful,
> + * but WITHOUT ANY WARRANTY; without even the implied warranty of
> + * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the GNU
> + * Lesser General Public License for more details.
> + *
> + * You should have received a copy of the GNU Lesser General Public
> + * License along with FFmpeg; if not, write to the Free Software
> + * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA
> + */
> +
> +#include "config.h"
> +#include "libavutil/attributes.h"
> +#include "libavutil/cpu.h"
> +#include "libavutil/x86/asm.h"
> +#include "libavutil/x86/cpu.h"
> +#include "libavcodec/apv_dsp.h"
> +
> +void ff_apv_decode_transquant_avx2(void *output,
> +                                   ptrdiff_t pitch,
> +                                   const int16_t *input,
> +                                   const int16_t *qmatrix,
> +                                   int64_t bit_depth,
> +                                   int64_t qp_shift);
> +
> +
> +av_cold void ff_apv_dsp_init_x86_64(APVDSPContext *dsp)
> +{
> +    int cpu_flags = av_get_cpu_flags();
> +
> +    if (EXTERNAL_AVX2_FAST(cpu_flags)) {
> +        dsp->decode_transquant = ff_apv_decode_transquant_avx2;
> +    }
> +}
> diff --git a/tests/checkasm/Makefile b/tests/checkasm/Makefile
> index d5c50e5599..193c1e4633 100644
> --- a/tests/checkasm/Makefile
> +++ b/tests/checkasm/Makefile
> @@ -28,6 +28,7 @@ AVCODECOBJS-$(CONFIG_AAC_DECODER)       += aacpsdsp.o \
>                                             sbrdsp.o
>  AVCODECOBJS-$(CONFIG_AAC_ENCODER)       += aacencdsp.o
>  AVCODECOBJS-$(CONFIG_ALAC_DECODER)      += alacdsp.o
> +AVCODECOBJS-$(CONFIG_APV_DECODER)       += apv_dsp.o
>  AVCODECOBJS-$(CONFIG_DCA_DECODER)       += synth_filter.o
>  AVCODECOBJS-$(CONFIG_DIRAC_DECODER)     += diracdsp.o
>  AVCODECOBJS-$(CONFIG_EXR_DECODER)       += exrdsp.o
> diff --git a/tests/checkasm/apv_dsp.c b/tests/checkasm/apv_dsp.c
> new file mode 100644
> index 0000000000..a0272d8edc
> --- /dev/null
> +++ b/tests/checkasm/apv_dsp.c
> @@ -0,0 +1,113 @@
> +/*
> + * This file is part of FFmpeg.
> + *
> + * FFmpeg is free software; you can redistribute it and/or
> + * modify it under the terms of the GNU Lesser General Public
> + * License as published by the Free Software Foundation; either
> + * version 2.1 of the License, or (at your option) any later version.
> + *
> + * FFmpeg is distributed in the hope that it will be useful,
> + * but WITHOUT ANY WARRANTY; without even the implied warranty of
> + * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the GNU
> + * Lesser General Public License for more details.
> + *
> + * You should have received a copy of the GNU Lesser General Public
> + * License along with FFmpeg; if not, write to the Free Software
> + * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA
> + */
> +
> +#include <stdint.h>
> +
> +#include "checkasm.h"
> +
> +#include "libavutil/attributes.h"
> +#include "libavutil/mem_internal.h"
> +#include "libavcodec/apv_dsp.h"
> +
> +
> +static void check_decode_transquant_8(void)
> +{
> +    LOCAL_ALIGNED_16(int16_t, input,      [64]);
> +    LOCAL_ALIGNED_16(int16_t, qmatrix,    [64]);
> +    LOCAL_ALIGNED_16(uint8_t, new_output, [64]);
> +    LOCAL_ALIGNED_16(uint8_t, ref_output, [64]);
> +
> +    declare_func(void,
> +                 uint8_t *output,
> +                 ptrdiff_t pitch,
> +                 const int16_t *input,
> +                 const int16_t *qmatrix,
> +                 int64_t bit_depth,
> +                 int64_t qp_shift);
> +
> +    for (int i = 0; i < 64; i++) {
> +        // Any signed 12-bit integer.
> +        input[i] = rnd() % 2048 - 1024;
> +
> +        // qmatrix input is premultiplied by level_scale, so
> +        // range is 1 to 255 * 71.  Interesting values are all
> +        // at the low end of that, though.
> +        qmatrix[i] = rnd() % 16 + 16;
> +    }
> +
> +    call_ref(ref_output, 8, input, qmatrix, 8, 4);
> +    call_new(new_output, 8, input, qmatrix, 8, 4);
> +
> +    for (int i = 0; i < 64; i++) {
> +        if (ref_output[i] != new_output[i])
> +            fail();
> +    }
> +
> +    bench_new(new_output, 8, input, qmatrix, 8, 4);
> +}
> +
> +static void check_decode_transquant_10(void)
> +{
> +    LOCAL_ALIGNED_16( int16_t, input,      [64]);
> +    LOCAL_ALIGNED_16( int16_t, qmatrix,    [64]);
> +    LOCAL_ALIGNED_16(uint16_t, new_output, [64]);
> +    LOCAL_ALIGNED_16(uint16_t, ref_output, [64]);
> +
> +    declare_func(void,
> +                 uint16_t *output,
> +                 ptrdiff_t pitch,
> +                 const int16_t *input,
> +                 const int16_t *qmatrix,
> +                 int64_t bit_depth,
> +                 int64_t qp_shift);
> +
> +    for (int i = 0; i < 64; i++) {
> +        // Any signed 14-bit integer.
> +        input[i] = rnd() % 16384 - 8192;
> +
> +        // qmatrix input is premultiplied by level_scale, so
> +        // range is 1 to 255 * 71.  Interesting values are all
> +        // at the low end of that, though.
> +        qmatrix[i] = 16; //rnd() % 16 + 16;
> +    }
> +
> +    call_ref(ref_output, 16, input, qmatrix, 10, 4);
> +    call_new(new_output, 16, input, qmatrix, 10, 4);
> +
> +    for (int i = 0; i < 64; i++) {
> +        if (ref_output[i] != new_output[i])
> +            fail();
> +    }
> +
> +    bench_new(new_output, 16, input, qmatrix, 10, 4);
> +}
> +
> +void checkasm_check_apv_dsp(void)
> +{
> +    APVDSPContext dsp;
> +
> +    ff_apv_dsp_init(&dsp);
> +
> +    if (check_func(dsp.decode_transquant, "decode_transquant_8"))
> +        check_decode_transquant_8();
> +
> +    if (check_func(dsp.decode_transquant, "decode_transquant_10"))
> +        check_decode_transquant_10();
> +
> +    report("apv_dsp");
> +}
> diff --git a/tests/checkasm/checkasm.c b/tests/checkasm/checkasm.c
> index 412b8b2cd1..3bb82ed0e5 100644
> --- a/tests/checkasm/checkasm.c
> +++ b/tests/checkasm/checkasm.c
> @@ -129,6 +129,9 @@ static const struct {
>      #if CONFIG_ALAC_DECODER
>          { "alacdsp", checkasm_check_alacdsp },
>      #endif
> +    #if CONFIG_APV_DECODER
> +        { "apv_dsp", checkasm_check_apv_dsp },
> +    #endif
>      #if CONFIG_AUDIODSP
>          { "audiodsp", checkasm_check_audiodsp },
>      #endif
> diff --git a/tests/checkasm/checkasm.h b/tests/checkasm/checkasm.h
> index ad239fb2a4..a6b5965e02 100644
> --- a/tests/checkasm/checkasm.h
> +++ b/tests/checkasm/checkasm.h
> @@ -83,6 +83,7 @@ void checkasm_check_ac3dsp(void);
>  void checkasm_check_aes(void);
>  void checkasm_check_afir(void);
>  void checkasm_check_alacdsp(void);
> +void checkasm_check_apv_dsp(void);
>  void checkasm_check_audiodsp(void);
>  void checkasm_check_av_tx(void);
>  void checkasm_check_blend(void);

_______________________________________________
ffmpeg-devel mailing list
ffmpeg-devel@ffmpeg.org
https://ffmpeg.org/mailman/listinfo/ffmpeg-devel

To unsubscribe, visit link above, or email
ffmpeg-devel-request@ffmpeg.org with subject "unsubscribe".