A lot of changes to replace ImageView and ImageViewMut structures on traits with the same names.

This commit is contained in:
Kirill Kuzminykh
2024-04-10 23:38:46 +03:00
parent 849d1f1601
commit d90d235fb2
110 changed files with 4250 additions and 3709 deletions
+83 -49
View File
@@ -1,3 +1,37 @@
## [Unreleased] - ReleaseDate
A lot of breaking changes have been done in this release:
- Structures `ImageView` and `ImageViewMut` have been removed. They always
did unnecessary memory allocation to store references to image rows.
Instead of these structures, the `ImageView` and `ImageViewMut` traits
have been added. The crate accepts any image container that provides
these traits.
- Also, traits `IntoImageView` and `IntoImageViewMut` have been added.
They allow you to write runtime adapters to convert your particular
image container into something that provides `ImageView`/`ImageViewMut` trait.
- `Resizer` now has two methods for resize:
- `resize()` accepts `IntoImageView` and `IntoImageViewMut` arguments;
- `resize_typed()` accepts `ImageView` and `ImageViewMut` arguments.
- Resize methods also accept the `options` argument.
With help of this argument, you can specify:
- how to crop the source image;
- whether to multiply the source image by the alpha channel and
divide the destination image by the alpha channel.
- The `MulDiv` implementation has been changed in the same way as `Resizer`.
It now has two versions of each method: dynamic and typed.
- Type of image dimensions has been changed from `NonZeroU32` into `u32`.
Now you can create and use images with zero pixels.
- Embedded implementation of image container `Image` moved from root of
the crate into module `images`.
- Added new image containers: `TypedImage`, `TypedImageMut`, `CroppedImage`
and `CroppedImageMut`.
- Added optional feature "image". It adds implementation of traits
`IntoImageView` and `IntoImageViewMut` for the
[DynamicImage](https://docs.rs/image/latest/image/enum.DynamicImage.html)
type from the `image` crate. It allows you to use `DynamicImage` instances
as arguments for `Resize::resize()` method.
## [3.0.4] - 2024-02-15
### Fixed
@@ -23,15 +57,15 @@
for `U8x3` and `U8x4` images.
- **BREAKING**: Changed internal data type for `U8x4` structure.
Now it is `[u8; 4]` instead of `u32`.
- Significantly improved (4.5 times on `x86_64`) speed of vertical convolution pass implemented
- Significantly improved (4.5 times on `x86_64`) speed of vertical convolution pass implemented
in native Rust for `U8`, `U8x2`, `U8x3` and `U8x4` images.
- Changed order of convolution passes for `U8`, `U8x2`, `U8x3` and `U8x4` images.
Now a vertical pass is the first and a horizontal pass is the second.
- **BREAKING**: Type of the `CropBox` fields has been changed to `f64`. Now you can use
fractional size and position of crop box.
- **BREAKING**: Type of the `centering` argument of `ImageView::set_crop_box_to_fit_dst_size()`
and `DynamicImageView::set_crop_box_to_fit_dst_size()` methods has been changed to `Optional<(f64, f64)>`.
- **BREAKING**: The `crop_box` argument of `ImageViewMut::crop()` and `DynamicImageViewMut::crop()`
and `DynamicImageView::set_crop_box_to_fit_dst_size()` methods has been changed to `Optional<(f64, f64)>`.
- **BREAKING**: The `crop_box` argument of `ImageViewMut::crop()` and `DynamicImageViewMut::crop()`
methods has been replaced with separate `left`, `top`, `width` and `height` arguments.
## [2.7.3] - 2023-05-07
@@ -46,7 +80,7 @@
### Fixed
- Added using of (read|write)_unaligned for unaligned pointers
- Added using of (read|write)_unaligned for unaligned pointers
on `arm64` and `wasm32` architectures.
([#15](https://github.com/Cykooz/fast_image_resize/issues/15)).
@@ -69,9 +103,9 @@
### Crate
- Slightly improved speed of `Convolution` implementation for `U8x2` images
- Slightly improved speed of `Convolution` implementation for `U8x2` images
and `Wasm32 SIMD128` instructions.
- Method `Image::buffer_mut()` was made public
- Method `Image::buffer_mut()` was made public
([#14](https://github.com/Cykooz/fast_image_resize/pull/14))
## [2.5.0] - 2023-01-29
@@ -94,7 +128,7 @@
- Slightly improved speed of `MulDiv` implementation for `U8x2`, `U8x4`, `U16x2` and `U16x4` images.
- Added optimisation for processing `U16x2` images by `MulDiv` with
helps of `NEON SIMD` instructions.
- Excluded possibility of unnecessary operations during resize
- Excluded possibility of unnecessary operations during resize
of cropped image by convolution algorithm.
- Added implementation `From` trait to convert `ImageViewMut` into `ImageView`.
- Added implementation `From` trait to convert `DynamicImageViewMut` into `DynamicImageView`.
@@ -139,25 +173,25 @@
### Crate
- Breaking changes:
- Struct `ImageView` replaced by enum `DynamicImageView`.
- Struct `ImageViewMut` replaced by enum `DynamicImageViewMut`.
- Trait `Pixel` renamed into `PixelExt` and some its internals changed:
- associated type `ComponentsCount` renamed into `CountOfComponents`.
- associated type `ComponentCountOfValues` deleted.
- associated method `components_count` renamed into `count_of_components`.
- associated method `component_count_of_values` renamed into `count_of_component_values`.
- All pixel types (`U8`, `U8x2`, ...) replaced by type aliases for new
generic structure `Pixel`. Use method `new()` to create
instance of one pixel.
- Struct `ImageView` replaced by enum `DynamicImageView`.
- Struct `ImageViewMut` replaced by enum `DynamicImageViewMut`.
- Trait `Pixel` renamed into `PixelExt` and some its internals changed:
- associated type `ComponentsCount` renamed into `CountOfComponents`.
- associated type `ComponentCountOfValues` deleted.
- associated method `components_count` renamed into `count_of_components`.
- associated method `component_count_of_values` renamed into `count_of_component_values`.
- All pixel types (`U8`, `U8x2`, ...) replaced by type aliases for new
generic structure `Pixel`. Use method `new()` to create
instance of one pixel.
- Added structure `PixelComponentMapper` that holds tables for mapping values of pixel's
components in forward and backward directions.
- Added function `create_gamma_22_mapper()` to create instance of `PixelComponentMapper`
that converts images with gamma 2.2 to linear colorspace and back.
that converts images with gamma 2.2 to linear colorspace and back.
- Added function `create_srgb_mapper()` to create instance of `PixelComponentMapper`
that converts images from SRGB colorspace to linear RGB and back.
- Added generic structs `ImageView` and `ImageViewMut`.
- Added functions `change_type_of_pixel_components` and
`change_type_of_pixel_components_dyn` that change type of pixel's
- Added functions `change_type_of_pixel_components` and
`change_type_of_pixel_components_dyn` that change type of pixel's
components in whole image.
- Added generic trait `IntoPixelComponent<Out: PixelComponent>`.
- Added generic structure `Pixel` for create all types of pixels.
@@ -168,7 +202,7 @@
### Example application
- Added option `--high_precision` to use `u16` as pixel components
- Added option `--high_precision` to use `u16` as pixel components
for intermediate image representation.
- Added converting of source image into linear colorspace before it will be resized.
Destination image will be returned into original colorspace before it will be saved.
@@ -179,14 +213,14 @@
## [0.9.7] - 2022-07-14
- Fixed resizing when the destination image has the same dimensions
- Fixed resizing when the destination image has the same dimensions
as the source image
([#9](https://github.com/Cykooz/fast_image_resize/issues/9)).
## [0.9.6] - 2022-06-28
- Added support of new type of pixels `PixelType::U16x4`.
- Fixed benchmarks for resizing images with alpha channel using
- Fixed benchmarks for resizing images with alpha channel using
the `resizer` crate.
- Removed `image` crate from benchmarks for resizing images with alpha.
- Added method `Image::copy(&self) -> Image<'static>`.
@@ -210,45 +244,45 @@
## [0.9.1] - 2022-05-12
- Added optimisation for processing `U8x2` images by `MulDiv` with
- Added optimisation for processing `U8x2` images by `MulDiv` with
helps of `SSE4.1` and `AVX2` instructions.
- Added optimisation for convolution of `U16x2` images with helps of
- Added optimisation for convolution of `U16x2` images with helps of
`AVX2` instructions.
## [0.9.0] - 2022-05-01
- Added support of new type of pixels `PixelType::U8x2`.
- Added into `MulDiv` support of images with pixel type `U8x2`.
- Added method `Image::into_vec(self) -> Vec<u8>`
- Added method `Image::into_vec(self) -> Vec<u8>`
([#7](https://github.com/Cykooz/fast_image_resize/pull/7)).
## [0.8.0] - 2022-03-23
- Added optimisation for convolution of U16x3 images with helps of `SSE4.1`
and `AVX2` instructions.
- Added partial optimisation for convolution of U8 images with helps of
- Added partial optimisation for convolution of U8 images with helps of
`SSE4.1` instructions.
- Allowed to create an instance of `Image`, `ImageVew` and `ImageViewMut`
from a buffer larger than necessary
- Allowed to create an instance of `Image`, `ImageVew` and `ImageViewMut`
from a buffer larger than necessary
([#5](https://github.com/Cykooz/fast_image_resize/issues/5)).
- Breaking changes:
- Removed methods: `Image::from_vec_u32()`, `Image::from_slice_u32()`.
- Removed error `InvalidBufferSizeError`.
- Removed methods: `Image::from_vec_u32()`, `Image::from_slice_u32()`.
- Removed error `InvalidBufferSizeError`.
## [0.7.0] - 2022-01-27
- Added support of new type of pixels `PixelType::U16x3`.
- Breaking changes:
- Added variant `U16x3` into the enum `PixelType`.
- Added variant `U16x3` into the enum `PixelType`.
## [0.6.0] - 2022-01-12
- Added optimisation of multiplying and dividing image by alpha channel with helps
of `SSE4.1` instructions.
- Improved performance of dividing image by alpha channel without forced
- Improved performance of dividing image by alpha channel without forced
SIMD instructions.
- Breaking changes:
- Deleted variant `SSE2` from enum `CpuExtensions`.
- Deleted variant `SSE2` from enum `CpuExtensions`.
## [0.5.3] - 2021-12-14
@@ -266,17 +300,17 @@
## [0.5.0] - 2021-11-18
- Added support of new type of pixels `PixelType::U8x3` (with
- Added support of new type of pixels `PixelType::U8x3` (with
auto-vectorization for SSE4.1).
- Exposed module `fast_image_resize::pixels` with types `U8x3`,
`U8x4`, `F32`, `I32`, `U8` used as wrappers for represent type of
- Exposed module `fast_image_resize::pixels` with types `U8x3`,
`U8x4`, `F32`, `I32`, `U8` used as wrappers for represent type of
one pixel of image.
- Some optimisations in code of convolution written in Rust (without
- Some optimisations in code of convolution written in Rust (without
intrinsics for SIMD).
- Breaking changes:
- Added variant `U8x3` into the enum `PixelType`.
- Changed internal tuple structures inside of variant of `ImageRows`
and `ImageRowsMut` enums.
- Added variant `U8x3` into the enum `PixelType`.
- Changed internal tuple structures inside of variant of `ImageRows`
and `ImageRowsMut` enums.
## [0.4.1] - 2021-11-13
@@ -286,10 +320,10 @@
- Added support of new type of pixels `PixelType::U8` (without forced SIMD).
- Breaking changes:
- `ImageData` renamed into `Image`.
- `SrcImageView` and `DstImageView` replaced by `ImageView`
and `ImageViewMut`.
- Method `Resizer.resize()` now returns `Result<(), DifferentTypesOfPixelsError>`.
- `ImageData` renamed into `Image`.
- `SrcImageView` and `DstImageView` replaced by `ImageView`
and `ImageViewMut`.
- Method `Resizer.resize()` now returns `Result<(), DifferentTypesOfPixelsError>`.
## [0.3.1] - 2021-10-09
@@ -299,10 +333,10 @@
- Added method `SrcImageView.set_crop_box_to_fit_dst_size()`.
- Fixed out-of-bounds error during resize with cropping.
- Refactored `ImageData`.
- Added methods: `from_vec_u32()`, `from_vec_u8()`, `from_slice_u32()`,
`from_slice_u8()`.
- Removed methods: `from_buffer()`, `from_pixels()`.
- Refactored `ImageData`.
- Added methods: `from_vec_u32()`, `from_vec_u8()`, `from_slice_u32()`,
`from_slice_u8()`.
- Removed methods: `from_buffer()`, `from_pixels()`.
## [0.2.0] - 2021-08-02
Generated
+654 -313
View File
File diff suppressed because it is too large Load Diff
+24 -18
View File
@@ -20,33 +20,39 @@ exclude = ["/data"]
[dependencies]
cfg-if = "1.0.0"
num-traits = "0.2.17"
thiserror = "1.0.56"
cfg-if = "1.0"
num-traits = "0.2.18"
thiserror = "1.0"
bytemuck = "1.15"
document-features = "0.2.8"
## Enable this feature to implement traits [IntoImageView](crate::IntoImageView) and
## [IntoImageViewMut](crate::IntoImageViewMut) for the
## [DynamicImage](https://docs.rs/image/latest/image/enum.DynamicImage.html)
## type from the `image` crate.
image = { version = "0.25.1", optional = true }
[features]
for_test = []
for_test = ["image"]
only_u8x4 = [] # This can be used to experiment with the crate's code.
[dev-dependencies]
fast_image_resize = { path = ".", features = ["for_test"] }
image = "0.24"
resize = "0.8"
rgb = "0.8"
png = "0.17"
resize = "0.8.4"
rgb = "0.8.37"
png = "0.17.13"
serde = { version = "1.0", features = ["serde_derive"] }
serde_json = "1"
walkdir = "2"
itertools = "0.12"
criterion = { version = "0.5", default-features = false, features = ["cargo_bench_support"] }
tera = "1"
serde_json = "1.0"
walkdir = "2.5"
itertools = "0.12.1"
criterion = { version = "0.5.1", default-features = false, features = ["cargo_bench_support"] }
tera = "1.19"
testing = { path = "testing" }
[target.'cfg(not(target_arch = "wasm32"))'.dev-dependencies]
nix = { version = "0.27", default-features = false, features = ["sched"] }
nix = { version = "0.28.0", default-features = false, features = ["sched"] }
[target.'cfg(all(not(target_arch = "wasm32"), not(target_os = "windows")))'.dev-dependencies]
@@ -115,14 +121,14 @@ debug = false
[profile.release]
opt-level = 3
#incremental = true
lto = true
incremental = true
#lto = true
#codegen-units = 1
strip = true
[profile.release.package.fast_image_resize]
codegen-units = 1
#[profile.release.package.fast_image_resize]
#codegen-units = 1
[profile.release.package.image]
+64 -99
View File
@@ -55,6 +55,7 @@ Other libraries used to compare of resizing speed:
- libvips (single-threaded mode, cache disabled)
<!-- bench_compare_rgb start -->
### Resize RGB8 image (U8x3) 4928x3279 => 852x567
Pipeline:
@@ -62,19 +63,21 @@ Pipeline:
`src_image => resize => dst_image`
- Source image [nasa-4928x3279.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279.png)
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:-----:|:--------:|:-------:|:--------:|
| image | 28.20 | - | 82.45 | 134.07 | 192.70 |
| resize | - | 26.83 | 53.56 | 97.73 | 144.63 |
| libvips | 7.73 | 60.66 | 19.84 | 30.15 | 39.46 |
| fir rust | 0.28 | 9.78 | 15.46 | 27.36 | 39.57 |
| fir sse4.1 | 0.28 | 3.87 | 5.59 | 9.89 | 15.44 |
| fir avx2 | 0.28 | 2.67 | 3.54 | 6.96 | 13.22 |
| image | 30.14 | - | 90.74 | 149.25 | 208.22 |
| resize | 7.78 | 26.82 | 53.54 | 97.38 | 144.44 |
| libvips | 7.78 | 59.56 | 18.69 | 30.36 | 39.69 |
| fir rust | 0.28 | 9.17 | 14.72 | 26.24 | 38.98 |
| fir sse4.1 | 0.28 | 4.08 | 5.79 | 10.32 | 15.94 |
| fir avx2 | 0.28 | 3.01 | 3.86 | 6.89 | 12.69 |
<!-- bench_compare_rgb end -->
<!-- bench_compare_rgba start -->
### Resize RGBA8 image (U8x4) 4928x3279 => 852x567
Pipeline:
@@ -83,19 +86,21 @@ Pipeline:
- Source image
[nasa-4928x3279-rgba.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279-rgba.png)
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
- The `image` crate does not support multiplying and dividing by alpha channel.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:------:|:--------:|:-------:|:--------:|
| resize | - | 42.96 | 85.43 | 147.79 | 211.49 |
| libvips | 10.06 | 122.80 | 188.97 | 338.18 | 499.99 |
| fir rust | 0.19 | 20.10 | 27.08 | 41.32 | 56.79 |
| fir sse4.1 | 0.19 | 10.03 | 12.24 | 18.57 | 25.15 |
| fir avx2 | 0.19 | 6.98 | 8.26 | 13.97 | 21.55 |
| resize | 11.30 | 42.85 | 85.27 | 147.28 | 211.34 |
| libvips | 9.15 | 120.11 | 188.46 | 337.77 | 499.37 |
| fir rust | 0.20 | 20.69 | 27.88 | 41.83 | 56.67 |
| fir sse4.1 | 0.19 | 10.19 | 12.43 | 17.99 | 24.64 |
| fir avx2 | 0.20 | 7.55 | 8.77 | 13.45 | 20.62 |
<!-- bench_compare_rgba end -->
<!-- bench_compare_l start -->
### Resize L8 image (U8) 4928x3279 => 852x567
Pipeline:
@@ -104,84 +109,67 @@ Pipeline:
- Source image [nasa-4928x3279.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279.png)
has converted into grayscale image with one byte per pixel.
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:-----:|:--------:|:-------:|:--------:|
| image | 25.96 | - | 56.78 | 84.17 | 112.12 |
| resize | - | 10.67 | 18.54 | 39.06 | 62.71 |
| libvips | 4.72 | 24.93 | 9.70 | 13.68 | 18.07 |
| fir rust | 0.15 | 4.08 | 5.24 | 7.48 | 11.33 |
| fir sse4.1 | 0.15 | 1.86 | 2.30 | 3.58 | 5.88 |
| fir avx2 | 0.15 | 1.66 | 1.86 | 2.24 | 4.21 |
| image | 27.05 | - | 58.63 | 86.87 | 115.57 |
| resize | 6.44 | 11.49 | 21.83 | 43.93 | 71.01 |
| libvips | 4.69 | 25.00 | 9.69 | 12.95 | 16.46 |
| fir rust | 0.15 | 3.97 | 4.98 | 7.15 | 11.04 |
| fir sse4.1 | 0.15 | 1.69 | 2.13 | 3.32 | 5.71 |
| fir avx2 | 0.15 | 1.73 | 1.94 | 2.30 | 4.33 |
<!-- bench_compare_l end -->
## Examples
### Resize RGBA8 image
Note:: You must enable `"image"` feature to support of
[image::DynamicImage](https://docs.rs/image/latest/image/enum.DynamicImage.html).
Otherwise, you have to convert such images into supported by the crate image type.
```rust
use std::io::BufWriter;
use std::num::NonZeroU32;
use image::codecs::png::PngEncoder;
use image::io::Reader as ImageReader;
use image::{ColorType, ImageEncoder};
use image::{ExtendedColorType, ImageEncoder};
use fast_image_resize as fr;
use fast_image_resize::{self as fr, IntoImageView};
fn main() {
// Read source image from file
let img = ImageReader::open("./data/nasa-4928x3279.png")
let src_image = ImageReader::open("./data/nasa-4928x3279.png")
.unwrap()
.decode()
.unwrap();
let width = NonZeroU32::new(img.width()).unwrap();
let height = NonZeroU32::new(img.height()).unwrap();
let mut src_image = fr::Image::from_vec_u8(
width,
height,
img.to_rgba8().into_raw(),
fr::PixelType::U8x4,
).unwrap();
// Multiple RGB channels of source image by alpha channel
// (not required for the Nearest algorithm)
let alpha_mul_div = fr::MulDiv::default();
alpha_mul_div
.multiply_alpha_inplace(&mut src_image.view_mut())
.unwrap();
// Create container for data of destination image
let dst_width = NonZeroU32::new(1024).unwrap();
let dst_height = NonZeroU32::new(768).unwrap();
let mut dst_image = fr::Image::new(
let dst_width = 1024;
let dst_height = 768;
let mut dst_image = fr::images::Image::new(
dst_width,
dst_height,
src_image.pixel_type(),
src_image.pixel_type().unwrap(),
);
// Get mutable view of destination image data
let mut dst_view = dst_image.view_mut();
// Create Resizer instance and resize source image
// into buffer of destination image
let mut resizer = fr::Resizer::new(
fr::ResizeAlg::Convolution(fr::FilterType::Lanczos3),
);
resizer.resize(&src_image.view(), &mut dst_view).unwrap();
// Divide RGB channels of destination image by alpha
alpha_mul_div.divide_alpha_inplace(&mut dst_view).unwrap();
resizer.resize(&src_image, &mut dst_image, None).unwrap();
// Write destination image as PNG-file
let mut result_buf = BufWriter::new(Vec::new());
PngEncoder::new(&mut result_buf)
.write_image(
dst_image.buffer(),
dst_width.get(),
dst_height.get(),
ColorType::Rgba8,
dst_width,
dst_height,
ExtendedColorType::Rgba8,
)
.unwrap();
}
@@ -190,63 +178,40 @@ fn main() {
### Resize with cropping
```rust
use std::num::NonZeroU32;
use image::codecs::png::PngEncoder;
use image::io::Reader as ImageReader;
use image::{ColorType, GenericImageView};
use fast_image_resize as fr;
fn resize_image_with_cropping(
mut src_view: fr::DynamicImageView,
dst_width: NonZeroU32,
dst_height: NonZeroU32
) -> fr::Image {
// Set cropping parameters
src_view.set_crop_box_to_fit_dst_size(
dst_width,
dst_height,
None,
);
// Create container for data of destination image
let mut dst_image = fr::Image::new(
dst_width,
dst_height,
src_view.pixel_type(),
);
// Get mutable view of destination image data
let mut dst_view = dst_image.view_mut();
// Create Resizer instance and resize source image
// into buffer of destination image
let mut resizer = fr::Resizer::new(
fr::ResizeAlg::Convolution(fr::FilterType::Lanczos3)
);
resizer.resize(&src_view, &mut dst_view).unwrap();
dst_image
}
use fast_image_resize::{self as fr, IntoImageView};
fn main() {
let img = ImageReader::open("./data/nasa-4928x3279.png")
.unwrap()
.decode()
.unwrap();
let width = NonZeroU32::new(img.width()).unwrap();
let height = NonZeroU32::new(img.height()).unwrap();
let src_image = fr::Image::from_vec_u8(
width,
height,
img.to_rgba8().into_raw(),
fr::PixelType::U8x4,
).unwrap();
resize_image_with_cropping(
src_image.view(),
NonZeroU32::new(1024).unwrap(),
NonZeroU32::new(768).unwrap(),
// Create container for data of destination image
let mut dst_image = fr::images::Image::new(
1024,
768,
img.pixel_type().unwrap(),
);
// Create Resizer instance and resize source image
// into buffer of destination image
let mut resizer = fr::Resizer::new(
fr::ResizeAlg::Convolution(fr::FilterType::Lanczos3),
);
resizer.resize(
&img,
&mut dst_image,
&fr::ResizeOptions::new().crop(
10.0, // left
10.0, // top
2000.0, // width
2000.0, // height
),
).unwrap();
}
```
+12 -24
View File
@@ -1,21 +1,15 @@
use std::num::NonZeroU32;
use fast_image_resize::images::Image;
use fast_image_resize::CpuExtensions;
use fast_image_resize::MulDiv;
use fast_image_resize::PixelType;
use fast_image_resize::{CpuExtensions, Image};
use testing::cpu_ext_into_str;
mod utils;
// Multiplies by alpha
fn get_src_image(
width: NonZeroU32,
height: NonZeroU32,
pixel_type: PixelType,
pixel: &[u8],
) -> Image<'static> {
let pixels_count = (width.get() * height.get()) as usize;
fn get_src_image(width: u32, height: u32, pixel_type: PixelType, pixel: &[u8]) -> Image<'static> {
let pixels_count = width as usize * height as usize;
let buffer = (0..pixels_count)
.flat_map(|_| pixel.iter().copied())
.collect();
@@ -28,8 +22,8 @@ fn multiplies_alpha(
cpu_extensions: CpuExtensions,
) {
let sample_size = 100;
let width = NonZeroU32::new(4096).unwrap();
let height = NonZeroU32::new(2048).unwrap();
let width = 4096;
let height = 2048;
let pixel: &[u8] = match pixel_type {
PixelType::U8x4 => &[255, 128, 0, 128],
PixelType::U8x2 => &[255, 128],
@@ -39,8 +33,6 @@ fn multiplies_alpha(
};
let src_data = get_src_image(width, height, pixel_type, pixel);
let mut dst_data = Image::new(width, height, pixel_type);
let src_view = src_data.view();
let mut dst_view = dst_data.view_mut();
let mut alpha_mul_div: MulDiv = Default::default();
unsafe {
alpha_mul_div.set_cpu_extensions(cpu_extensions);
@@ -54,7 +46,7 @@ fn multiplies_alpha(
|bencher| {
bencher.iter(|| {
alpha_mul_div
.multiply_alpha(&src_view, &mut dst_view)
.multiply_alpha(&src_data, &mut dst_data)
.unwrap();
})
},
@@ -68,9 +60,8 @@ fn multiplies_alpha(
cpu_ext_into_str(cpu_extensions),
|bencher| {
let mut image = src_image.copy();
let mut view = image.view_mut();
bencher.iter(|| {
alpha_mul_div.multiply_alpha_inplace(&mut view).unwrap();
alpha_mul_div.multiply_alpha_inplace(&mut image).unwrap();
})
},
);
@@ -82,8 +73,8 @@ fn divides_alpha(
cpu_extensions: CpuExtensions,
) {
let sample_size = 100;
let width = NonZeroU32::new(4095).unwrap();
let height = NonZeroU32::new(2048).unwrap();
let width = 4095;
let height = 2048;
let pixel: &[u8] = match pixel_type {
PixelType::U8x4 => &[128, 64, 0, 128],
PixelType::U8x2 => &[128, 128],
@@ -93,8 +84,6 @@ fn divides_alpha(
};
let src_data = get_src_image(width, height, pixel_type, pixel);
let mut dst_data = Image::new(width, height, pixel_type);
let src_view = src_data.view();
let mut dst_view = dst_data.view_mut();
let mut alpha_mul_div: MulDiv = Default::default();
unsafe {
alpha_mul_div.set_cpu_extensions(cpu_extensions);
@@ -108,7 +97,7 @@ fn divides_alpha(
|bencher| {
bencher.iter(|| {
alpha_mul_div
.divide_alpha(&src_view, &mut dst_view)
.divide_alpha(&src_data, &mut dst_data)
.unwrap();
})
},
@@ -122,9 +111,8 @@ fn divides_alpha(
cpu_ext_into_str(cpu_extensions),
|bencher| {
let mut image = src_image.copy();
let mut view = image.view_mut();
bencher.iter(|| {
alpha_mul_div.divide_alpha_inplace(&mut view).unwrap();
alpha_mul_div.divide_alpha_inplace(&mut image).unwrap();
})
},
);
+3 -4
View File
@@ -1,5 +1,6 @@
use fast_image_resize::create_srgb_mapper;
use fast_image_resize::images::Image;
use fast_image_resize::pixels::U8x3;
use fast_image_resize::{create_srgb_mapper, Image};
use testing::PixelTestingExt;
mod utils;
@@ -11,12 +12,10 @@ pub fn bench_color_mapper(bench_group: &mut utils::BenchGroup) {
src_image.height(),
src_image.pixel_type(),
);
let src_view = src_image.view();
let mut dst_view = dst_image.view_mut();
let mapper = create_srgb_mapper();
bench_group.bench_function("SRGB U8x3 => RGB U8x3", |bencher| {
bencher.iter(|| {
mapper.forward_map(&src_view, &mut dst_view).unwrap();
mapper.forward_map(&src_image, &mut dst_image).unwrap();
})
});
}
+1 -1
View File
@@ -18,7 +18,7 @@ pub fn bench_compare_l(bench_group: &mut utils::BenchGroup) {
src_image.height(),
);
utils::libvips_resize::<P>(bench_group, false);
utils::fir_resize::<P>(bench_group);
utils::fir_resize::<P>(bench_group, false);
}
fn main() {
+1 -1
View File
@@ -18,7 +18,7 @@ pub fn bench_downscale_l16(bench_group: &mut utils::BenchGroup) {
src_image.height(),
);
utils::libvips_resize::<P>(bench_group, false);
utils::fir_resize::<P>(bench_group);
utils::fir_resize::<P>(bench_group, false);
}
fn main() {
+1 -1
View File
@@ -5,7 +5,7 @@ mod utils;
pub fn bench_downscale_la(bench_group: &mut utils::BenchGroup) {
type P = U8x2;
utils::libvips_resize::<P>(bench_group, true);
utils::fir_resize_with_alpha::<P>(bench_group);
utils::fir_resize::<P>(bench_group, true);
}
fn main() {
+1 -1
View File
@@ -5,7 +5,7 @@ mod utils;
pub fn bench_downscale_la16(bench_group: &mut utils::BenchGroup) {
type P = U16x2;
utils::libvips_resize::<P>(bench_group, true);
utils::fir_resize_with_alpha::<P>(bench_group);
utils::fir_resize::<P>(bench_group, true);
}
fn main() {
+1 -1
View File
@@ -18,7 +18,7 @@ pub fn bench_downscale_rgb(bench_group: &mut utils::BenchGroup) {
src_image.height(),
);
utils::libvips_resize::<P>(bench_group, false);
utils::fir_resize::<P>(bench_group);
utils::fir_resize::<P>(bench_group, false);
}
fn main() {
+1 -1
View File
@@ -18,7 +18,7 @@ pub fn bench_downscale_rgb16(bench_group: &mut utils::BenchGroup) {
src_image.height(),
);
utils::libvips_resize::<P>(bench_group, false);
utils::fir_resize::<P>(bench_group);
utils::fir_resize::<P>(bench_group, false);
}
fn main() {
+1 -1
View File
@@ -17,7 +17,7 @@ pub fn bench_downscale_rgba(bench_group: &mut utils::BenchGroup) {
src_image.height(),
);
utils::libvips_resize::<P>(bench_group, true);
utils::fir_resize_with_alpha::<P>(bench_group);
utils::fir_resize::<P>(bench_group, true);
}
fn main() {
+1 -1
View File
@@ -17,7 +17,7 @@ pub fn bench_downscale_rgba16(bench_group: &mut utils::BenchGroup) {
src_image.height(),
);
utils::libvips_resize::<P>(bench_group, true);
utils::fir_resize_with_alpha::<P>(bench_group);
utils::fir_resize::<P>(bench_group, true);
}
fn main() {
+20 -32
View File
@@ -1,42 +1,35 @@
use fast_image_resize::images::Image;
use fast_image_resize::pixels::*;
use fast_image_resize::Image;
use fast_image_resize::ResizeOptions;
use fast_image_resize::{CpuExtensions, FilterType, PixelType, ResizeAlg, Resizer};
use std::num::NonZeroU32;
use testing::{cpu_ext_into_str, nonzero, PixelTestingExt};
use testing::{cpu_ext_into_str, PixelTestingExt};
mod utils;
const NEW_SIZE: u32 = 695;
fn native_nearest_u8x4_bench(bench_group: &mut utils::BenchGroup) {
let image = U8x4::load_big_square_src_image();
let mut res_image = Image::new(nonzero(NEW_SIZE), nonzero(NEW_SIZE), image.pixel_type());
let src_image = image.view();
let mut dst_image = res_image.view_mut();
let src_image = U8x4::load_big_square_src_image();
let mut dst_image = Image::new(NEW_SIZE, NEW_SIZE, PixelType::U8x4);
let mut resizer = Resizer::new(ResizeAlg::Nearest);
unsafe {
resizer.set_cpu_extensions(CpuExtensions::None);
}
utils::bench(bench_group, 100, "U8x4 Nearest", "rust", |bencher| {
bencher.iter(|| {
resizer.resize(&src_image, &mut dst_image).unwrap();
})
bencher.iter(|| resizer.resize(&src_image, &mut dst_image, None).unwrap())
});
}
#[cfg(not(feature = "only_u8x4"))]
fn native_nearest_u8_bench(bench_group: &mut utils::BenchGroup) {
let image = U8::load_big_square_src_image();
let mut res_image = Image::new(nonzero(NEW_SIZE), nonzero(NEW_SIZE), image.pixel_type());
let src_image = image.view();
let mut dst_image = res_image.view_mut();
let src_image = U8::load_big_square_src_image();
let mut dst_image = Image::new(NEW_SIZE, NEW_SIZE, PixelType::U8);
let mut resizer = Resizer::new(ResizeAlg::Nearest);
unsafe {
resizer.set_cpu_extensions(CpuExtensions::None);
}
utils::bench(bench_group, 100, "U8 Nearest", "rust", |bencher| {
bencher.iter(|| {
resizer.resize(&src_image, &mut dst_image).unwrap();
})
bencher.iter(|| resizer.resize(&src_image, &mut dst_image, None).unwrap())
});
}
@@ -45,14 +38,13 @@ fn downscale_bench(
image: &Image<'static>,
cpu_extensions: CpuExtensions,
filter_type: FilterType,
dst_width: NonZeroU32,
dst_height: NonZeroU32,
dst_width: u32,
dst_height: u32,
name_prefix: &str,
) {
let mut res_image = Image::new(dst_width, dst_height, image.pixel_type());
let src_image = image.view();
let mut dst_image = res_image.view_mut();
let mut resizer = Resizer::new(ResizeAlg::Convolution(filter_type));
let options = ResizeOptions::new().use_alpha(false);
unsafe {
resizer.set_cpu_extensions(cpu_extensions);
}
@@ -66,11 +58,7 @@ fn downscale_bench(
100,
&format!("{:?} {:?}", image.pixel_type(), filter_type),
&format!("{}{}", cpu_ext_into_str(cpu_extensions), prefix),
|bencher| {
bencher.iter(|| {
resizer.resize(&src_image, &mut dst_image).unwrap();
})
},
|bencher| bencher.iter(|| resizer.resize(image, &mut res_image, &options).unwrap()),
);
}
@@ -122,7 +110,7 @@ pub fn resize_in_one_dimension_bench(bench_group: &mut utils::BenchGroup) {
&image,
cpu_extension,
FilterType::Lanczos3,
nonzero(NEW_SIZE),
NEW_SIZE,
image.height(),
"H",
);
@@ -132,7 +120,7 @@ pub fn resize_in_one_dimension_bench(bench_group: &mut utils::BenchGroup) {
cpu_extension,
FilterType::Lanczos3,
image.height(),
nonzero(NEW_SIZE),
NEW_SIZE,
"V",
);
}
@@ -188,8 +176,8 @@ pub fn resize_bench(bench_group: &mut utils::BenchGroup) {
&image,
cpu_extension,
FilterType::Lanczos3,
nonzero(NEW_SIZE),
nonzero(NEW_SIZE),
NEW_SIZE,
NEW_SIZE,
"",
);
}
@@ -200,12 +188,12 @@ pub fn resize_bench(bench_group: &mut utils::BenchGroup) {
native_nearest_u8_bench(bench_group);
}
fn main1() {
fn main() {
let results = utils::run_bench(resize_bench, "Resize");
println!("{}", utils::build_md_table(&results));
}
fn main() {
fn main2() {
let results = utils::run_bench(resize_in_one_dimension_bench, "Resize one dimension");
println!("{}", utils::build_md_table(&results));
}
+6 -6
View File
@@ -8,18 +8,18 @@ Environment:
- CPU: AMD Ryzen 9 5950X
- RAM: DDR4 3800 MHz
{% endif -%}
- Ubuntu 22.04 (linux 6.2.0)
- Rust 1.75.0
- Ubuntu 22.04 (linux 6.5.0)
- Rust 1.77.1
- criterion = "0.5.1"
- fast_image_resize = "3.0.0"
- fast_image_resize = "4.0.0"
{% if arch_id == "wasm32" -%}
- wasmtime = "16.0.0"
- wasmtime = "19.0.1"
{% endif %}
Other libraries used to compare of resizing speed:
- image = "0.24.7" (<https://crates.io/crates/image>)
- resize = "0.8.3" (<https://crates.io/crates/resize>)
- image = "0.25.1" (<https://crates.io/crates/image>)
- resize = "0.8.4" (<https://crates.io/crates/resize>)
{% if arch_id != "wasm32" -%}
- libvips = "8.12.1" (single-threaded mode, cache disabled)
{% endif %}
+13 -102
View File
@@ -3,8 +3,9 @@ use std::ops::Deref;
use criterion::black_box;
use image::{imageops, ImageBuffer};
use fast_image_resize::{CpuExtensions, FilterType, Image, MulDiv, ResizeAlg, Resizer};
use testing::{cpu_ext_into_str, nonzero, PixelTestingExt};
use fast_image_resize::images::Image;
use fast_image_resize::{CpuExtensions, FilterType, ResizeAlg, ResizeOptions, Resizer};
use testing::{cpu_ext_into_str, PixelTestingExt};
use crate::utils::bencher::{bench, BenchGroup};
@@ -45,23 +46,15 @@ pub fn resize_resize<Format, Out>(
Out: Clone,
Format: resize::PixelFormat<OutputPixel = Out> + Copy,
{
fn box_kernel(_: f32) -> f32 {
1.0
}
for alg_name in ALG_NAMES {
if alg_name == "Nearest" {
// "resize" doesn't support "nearest" algorithm
continue;
}
let mut dst =
vec![pixel_format.into_pixel(Format::new()); (NEW_WIDTH * NEW_HEIGHT) as usize];
let sample_size = if alg_name == "Lanczos3" { 60 } else { 100 };
bench(bench_group, sample_size, "resize", alg_name, |bencher| {
let filter = match alg_name {
"Box" => resize::Type::Custom(resize::Filter::new(Box::new(box_kernel), 0.5)),
"Nearest" => resize::Type::Point,
"Box" => resize::Type::Custom(resize::Filter::box_filter(0.5)),
"Bilinear" => resize::Type::Triangle,
"Bicubic" => resize::Type::Catrom,
"Lanczos3" => resize::Type::Lanczos3,
@@ -109,8 +102,8 @@ mod vips {
app.cache_set_max_mem(0);
let src_image_data = P::load_big_src_image();
let src_width = src_image_data.width().get() as i32;
let src_height = src_image_data.height().get() as i32;
let src_width = src_image_data.width() as i32;
let src_height = src_image_data.height() as i32;
let band_format = if P::count_of_component_values() > 256 {
BandFormat::Ushort
} else {
@@ -186,16 +179,10 @@ mod vips {
}
/// Resize image with help of "fast_imager_resize" crate
pub fn fir_resize<P: PixelTestingExt>(bench_group: &mut BenchGroup) {
pub fn fir_resize<P: PixelTestingExt>(bench_group: &mut BenchGroup, use_alpha: bool) {
let resize_options = ResizeOptions::new().use_alpha(use_alpha);
let src_image_data = P::load_big_src_image();
let src_view = src_image_data.view();
let mut dst_image = Image::new(
nonzero(NEW_WIDTH),
nonzero(NEW_HEIGHT),
src_view.pixel_type(),
);
let mut dst_view = dst_image.view_mut();
let mut dst_image = Image::new(NEW_WIDTH, NEW_HEIGHT, src_image_data.pixel_type());
let mut cpu_extensions = vec![CpuExtensions::None];
#[cfg(target_arch = "x86_64")]
{
@@ -232,89 +219,13 @@ pub fn fir_resize<P: PixelTestingExt>(bench_group: &mut BenchGroup) {
format!("fir {}", cpu_ext_into_str(cpu_ext)),
alg_name,
|bencher| {
fast_resizer.reset_internal_buffers();
bencher.iter(|| {
fast_resizer.resize(&src_view, &mut dst_view).unwrap();
fast_resizer
.resize(&src_image_data, &mut dst_image, &resize_options)
.unwrap()
})
},
);
}
}
}
/// Resize image with alpha channel with help of "fast_imager_resize" crate
pub fn fir_resize_with_alpha<P: PixelTestingExt>(bench_group: &mut BenchGroup) {
let src_image = P::load_big_src_image();
let src_view = src_image.view();
let mut premultiplied_src_image =
Image::new(src_image.width(), src_image.height(), src_view.pixel_type());
let mut dst_image = Image::new(
nonzero(NEW_WIDTH),
nonzero(NEW_HEIGHT),
src_view.pixel_type(),
);
let mut dst_view = dst_image.view_mut();
let mut mul_div = MulDiv::default();
let mut cpu_ext_and_name = vec![CpuExtensions::None];
#[cfg(target_arch = "x86_64")]
{
cpu_ext_and_name.push(CpuExtensions::Sse4_1);
cpu_ext_and_name.push(CpuExtensions::Avx2);
}
#[cfg(target_arch = "aarch64")]
{
cpu_ext_and_name.push(CpuExtensions::Neon);
}
#[cfg(target_arch = "wasm32")]
{
cpu_ext_and_name.push(CpuExtensions::Simd128);
}
for cpu_ext in cpu_ext_and_name {
for alg_name in ALG_NAMES {
let resize_alg = match alg_name {
"Nearest" => ResizeAlg::Nearest,
"Box" => ResizeAlg::Convolution(FilterType::Box),
"Bilinear" => ResizeAlg::Convolution(FilterType::Bilinear),
"Bicubic" => ResizeAlg::Convolution(FilterType::CatmullRom),
"Lanczos3" => ResizeAlg::Convolution(FilterType::Lanczos3),
_ => return,
};
let mut fast_resizer = Resizer::new(resize_alg);
unsafe {
fast_resizer.set_cpu_extensions(cpu_ext);
mul_div.set_cpu_extensions(cpu_ext);
}
let sample_size = 100;
bench(
bench_group,
sample_size,
format!("fir {}", cpu_ext_into_str(cpu_ext)),
alg_name,
|bencher| {
fast_resizer.reset_internal_buffers();
match resize_alg {
ResizeAlg::Nearest => {
bencher.iter(|| {
fast_resizer.resize(&src_view, &mut dst_view).unwrap();
});
}
_ => {
bencher.iter(|| {
let mut premultiplied_view = premultiplied_src_image.view_mut();
mul_div
.multiply_alpha(&src_view, &mut premultiplied_view)
.unwrap();
fast_resizer
.resize(&premultiplied_view.into(), &mut dst_view)
.unwrap();
mul_div.divide_alpha_inplace(&mut dst_view).unwrap();
});
}
};
},
);
}
}
}
+50 -50
View File
@@ -35,16 +35,16 @@ Pipeline:
`src_image => resize => dst_image`
- Source image [nasa-4928x3279.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279.png)
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:-----:|:--------:|:-------:|:--------:|
| image | 28.20 | - | 82.45 | 134.07 | 192.70 |
| resize | - | 26.83 | 53.56 | 97.73 | 144.63 |
| libvips | 7.73 | 60.66 | 19.84 | 30.15 | 39.46 |
| fir rust | 0.28 | 9.78 | 15.46 | 27.36 | 39.57 |
| fir sse4.1 | 0.28 | 3.87 | 5.59 | 9.89 | 15.44 |
| fir avx2 | 0.28 | 2.67 | 3.54 | 6.96 | 13.22 |
| image | 30.14 | - | 90.74 | 149.25 | 208.22 |
| resize | 7.78 | 26.82 | 53.54 | 97.38 | 144.44 |
| libvips | 7.78 | 59.56 | 18.69 | 30.36 | 39.69 |
| fir rust | 0.28 | 9.17 | 14.72 | 26.24 | 38.98 |
| fir sse4.1 | 0.28 | 4.08 | 5.79 | 10.32 | 15.94 |
| fir avx2 | 0.28 | 3.01 | 3.86 | 6.89 | 12.69 |
<!-- bench_compare_rgb end -->
<!-- bench_compare_rgba start -->
@@ -56,16 +56,16 @@ Pipeline:
- Source image
[nasa-4928x3279-rgba.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279-rgba.png)
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
- The `image` crate does not support multiplying and dividing by alpha channel.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:------:|:--------:|:-------:|:--------:|
| resize | - | 42.96 | 85.43 | 147.79 | 211.49 |
| libvips | 10.06 | 122.80 | 188.97 | 338.18 | 499.99 |
| fir rust | 0.19 | 20.10 | 27.08 | 41.32 | 56.79 |
| fir sse4.1 | 0.19 | 10.03 | 12.24 | 18.57 | 25.15 |
| fir avx2 | 0.19 | 6.98 | 8.26 | 13.97 | 21.55 |
| resize | 11.30 | 42.85 | 85.27 | 147.28 | 211.34 |
| libvips | 9.15 | 120.11 | 188.46 | 337.77 | 499.37 |
| fir rust | 0.20 | 20.69 | 27.88 | 41.83 | 56.67 |
| fir sse4.1 | 0.19 | 10.19 | 12.43 | 17.99 | 24.64 |
| fir avx2 | 0.20 | 7.55 | 8.77 | 13.45 | 20.62 |
<!-- bench_compare_rgba end -->
<!-- bench_compare_l start -->
@@ -77,16 +77,16 @@ Pipeline:
- Source image [nasa-4928x3279.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279.png)
has converted into grayscale image with one byte per pixel.
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:-----:|:--------:|:-------:|:--------:|
| image | 25.96 | - | 56.78 | 84.17 | 112.12 |
| resize | - | 10.67 | 18.54 | 39.06 | 62.71 |
| libvips | 4.72 | 24.93 | 9.70 | 13.68 | 18.07 |
| fir rust | 0.15 | 4.08 | 5.24 | 7.48 | 11.33 |
| fir sse4.1 | 0.15 | 1.86 | 2.30 | 3.58 | 5.88 |
| fir avx2 | 0.15 | 1.66 | 1.86 | 2.24 | 4.21 |
| image | 27.05 | - | 58.63 | 86.87 | 115.57 |
| resize | 6.44 | 11.49 | 21.83 | 43.93 | 71.01 |
| libvips | 4.69 | 25.00 | 9.69 | 12.95 | 16.46 |
| fir rust | 0.15 | 3.97 | 4.98 | 7.15 | 11.04 |
| fir sse4.1 | 0.15 | 1.69 | 2.13 | 3.32 | 5.71 |
| fir avx2 | 0.15 | 1.73 | 1.94 | 2.30 | 4.33 |
<!-- bench_compare_l end -->
<!-- bench_compare_la start -->
@@ -99,16 +99,16 @@ Pipeline:
- Source image
[nasa-4928x3279-rgba.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279-rgba.png)
has converted into grayscale image with alpha channel (two bytes per pixel).
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
- The `image` crate does not support multiplying and dividing by alpha channel.
- The `resize` crate does not support this pixel format.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:-----:|:--------:|:-------:|:--------:|
| libvips | 6.48 | 73.12 | 117.76 | 207.96 | 293.16 |
| fir rust | 0.17 | 11.19 | 12.90 | 17.42 | 23.90 |
| fir sse4.1 | 0.17 | 6.16 | 7.21 | 9.74 | 13.56 |
| fir avx2 | 0.17 | 3.95 | 4.57 | 6.41 | 9.24 |
| libvips | 6.54 | 73.82 | 118.25 | 206.24 | 293.99 |
| fir rust | 0.18 | 10.95 | 12.78 | 17.17 | 23.78 |
| fir sse4.1 | 0.17 | 6.21 | 7.31 | 9.73 | 14.07 |
| fir avx2 | 0.17 | 4.23 | 4.72 | 6.26 | 8.87 |
<!-- bench_compare_la end -->
<!-- bench_compare_rgb16 start -->
@@ -120,16 +120,16 @@ Pipeline:
- Source image [nasa-4928x3279.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279.png)
has converted into RGB16 image.
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:-----:|:--------:|:-------:|:--------:|
| image | 28.92 | - | 82.94 | 134.72 | 185.59 |
| resize | - | 26.91 | 49.69 | 95.90 | 141.39 |
| libvips | 16.00 | 63.29 | 54.38 | 102.70 | 126.07 |
| fir rust | 0.34 | 26.13 | 42.62 | 77.16 | 112.71 |
| fir sse4.1 | 0.34 | 16.06 | 23.04 | 36.76 | 51.99 |
| fir avx2 | 0.34 | 13.99 | 19.70 | 30.89 | 38.32 |
| image | 30.85 | - | 84.62 | 136.64 | 190.22 |
| resize | 8.09 | 26.36 | 50.24 | 96.74 | 143.86 |
| libvips | 16.01 | 62.96 | 54.26 | 102.92 | 125.23 |
| fir rust | 0.34 | 25.73 | 39.52 | 67.32 | 96.71 |
| fir sse4.1 | 0.36 | 16.25 | 23.15 | 36.67 | 51.82 |
| fir avx2 | 0.34 | 14.02 | 19.32 | 30.01 | 37.16 |
<!-- bench_compare_rgb16 end -->
<!-- bench_compare_rgba16 start -->
@@ -141,16 +141,16 @@ Pipeline:
- Source image
[nasa-4928x3279-rgba.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279-rgba.png)
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
- The `image` crate does not support multiplying and dividing by alpha channel.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:------:|:--------:|:-------:|:--------:|
| resize | - | 43.62 | 84.16 | 144.69 | 207.01 |
| libvips | 22.70 | 130.12 | 205.66 | 365.03 | 536.16 |
| fir rust | 0.38 | 60.71 | 79.18 | 116.62 | 155.54 |
| fir sse4.1 | 0.38 | 32.14 | 42.66 | 64.57 | 86.63 |
| fir avx2 | 0.38 | 20.34 | 25.74 | 36.85 | 48.39 |
| resize | 11.89 | 43.39 | 83.74 | 144.33 | 206.71 |
| libvips | 22.60 | 129.23 | 205.57 | 367.46 | 538.41 |
| fir rust | 0.37 | 58.67 | 77.53 | 114.73 | 153.99 |
| fir sse4.1 | 0.38 | 31.85 | 42.51 | 63.89 | 86.03 |
| fir avx2 | 0.39 | 20.00 | 25.44 | 36.04 | 47.38 |
<!-- bench_compare_rgba16 end -->
<!-- bench_compare_l16 start -->
@@ -162,16 +162,16 @@ Pipeline:
- Source image [nasa-4928x3279.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279.png)
has converted into grayscale image with two bytes per pixel.
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:-----:|:--------:|:-------:|:--------:|
| image | 26.00 | - | 57.17 | 85.75 | 114.97 |
| resize | - | 9.95 | 16.13 | 33.72 | 58.90 |
| libvips | 7.87 | 26.35 | 21.70 | 36.52 | 45.96 |
| fir rust | 0.17 | 14.11 | 20.97 | 29.53 | 40.03 |
| fir sse4.1 | 0.17 | 5.59 | 7.71 | 13.11 | 19.02 |
| fir avx2 | 0.17 | 5.70 | 6.67 | 8.79 | 13.84 |
| image | 27.48 | - | 59.32 | 88.21 | 117.58 |
| resize | 6.29 | 11.31 | 20.55 | 44.20 | 68.86 |
| libvips | 7.89 | 26.35 | 20.67 | 36.39 | 46.13 |
| fir rust | 0.17 | 13.64 | 19.76 | 28.76 | 39.82 |
| fir sse4.1 | 0.17 | 5.34 | 7.53 | 12.93 | 18.85 |
| fir avx2 | 0.17 | 5.45 | 6.35 | 8.45 | 13.52 |
<!-- bench_compare_l16 end -->
<!-- bench_compare_la16 start -->
@@ -184,14 +184,14 @@ Pipeline:
- Source image
[nasa-4928x3279-rgba.png](https://github.com/Cykooz/fast_image_resize/blob/main/data/nasa-4928x3279-rgba.png)
has converted into grayscale image with alpha channel (four bytes per pixel).
- Numbers in table is mean duration of image resizing in milliseconds.
- Numbers in table are mean duration of image resizing in milliseconds.
- The `image` crate does not support multiplying and dividing by alpha channel.
- The `resize` crate does not support this pixel format.
| | Nearest | Box | Bilinear | Bicubic | Lanczos3 |
|------------|:-------:|:-----:|:--------:|:-------:|:--------:|
| libvips | 12.55 | 79.43 | 133.92 | 232.67 | 328.01 |
| fir rust | 0.19 | 27.70 | 36.68 | 51.82 | 72.60 |
| fir sse4.1 | 0.19 | 15.27 | 21.39 | 33.56 | 46.19 |
| fir avx2 | 0.19 | 11.53 | 14.72 | 21.77 | 28.98 |
| libvips | 12.50 | 79.75 | 133.81 | 232.26 | 329.04 |
| fir rust | 0.20 | 25.00 | 32.91 | 51.64 | 71.49 |
| fir sse4.1 | 0.20 | 15.02 | 21.36 | 34.09 | 46.03 |
| fir avx2 | 0.20 | 11.62 | 14.91 | 21.95 | 29.14 |
<!-- bench_compare_la16 end -->
+7 -7
View File
@@ -5,14 +5,14 @@ edition = "2021"
[dependencies]
fast_image_resize = {path=".."}
image = "0.24"
clap = { version = "4", features = ["derive"] }
log = "0.4"
env_logger = "0.10"
fast_image_resize = { path = "..", features = ["image"] }
image = "0.25.1"
clap = { version = "4.5", features = ["derive"] }
log = "0.4.21"
env_logger = "0.11.3"
anyhow = "1.0"
clap-verbosity-flag = "2"
once_cell = "1"
clap-verbosity-flag = "2.2"
once_cell = "1.19"
[package.metadata.release]
+35 -62
View File
@@ -1,5 +1,4 @@
use std::ffi::OsStr;
use std::num::NonZeroU32;
use std::path::PathBuf;
use anyhow::{anyhow, Context, Result};
@@ -10,13 +9,19 @@ use log::debug;
use once_cell::sync::Lazy;
use fast_image_resize as fr;
use fast_image_resize::images::Image;
use fast_image_resize::ResizeOptions;
mod structs;
#[derive(Parser)]
#[clap(author = "Kirill K.")]
#[clap(version, about, long_about = None)]
#[clap(disable_help_flag = true)]
struct Cli {
#[clap(long, action = clap::ArgAction::HelpLong)]
help: Option<bool>,
/// Path to source image file
#[clap(value_parser)]
source_path: PathBuf,
@@ -69,40 +74,25 @@ fn main() -> Result<()> {
}
fn resize(cli: &Cli) -> Result<()> {
let (mut src_image, color_type, orig_pixel_type) = open_source_image(cli)?;
let (src_image, color_type, orig_pixel_type) = open_source_image(cli)?;
let mut dst_image = create_destination_image(cli, &src_image);
let mul_div = fr::MulDiv::default();
let algorithm = get_resizing_algorithm(cli);
let mut resizer = fr::Resizer::new(algorithm);
if color_type.has_alpha() {
debug!("Multiply color channels of the source image by alpha channel");
mul_div
.multiply_alpha_inplace(&mut src_image.view_mut())
.with_context(|| "Failed to multiply color channels by alpha")?;
}
debug!(
"Resize the source image into {}x{}",
dst_image.width(),
dst_image.height()
);
resizer
.resize(&src_image.view(), &mut dst_image.view_mut())
.resize(&src_image, &mut dst_image, None)
.with_context(|| "Failed to resize image")?;
if color_type.has_alpha() {
debug!("Divide color channels of the result image by alpha channel");
mul_div
.divide_alpha_inplace(&mut dst_image.view_mut())
.with_context(|| "Failed to divide color channels by alpha")?;
}
save_result(cli, dst_image, color_type, orig_pixel_type)
}
fn open_source_image(cli: &Cli) -> Result<(fr::Image<'static>, ColorType, fr::PixelType)> {
fn open_source_image(cli: &Cli) -> Result<(Image<'static>, ColorType, fr::PixelType)> {
let source_path = &cli.source_path;
debug!("Opening the source image {:?}", source_path);
let image = ImageReader::open(source_path)
@@ -110,12 +100,10 @@ fn open_source_image(cli: &Cli) -> Result<(fr::Image<'static>, ColorType, fr::Pi
.decode()
.with_context(|| "Failed to decode source image")?;
let src_width = NonZeroU32::new(image.width())
.with_context(|| "Failed to get width of the source image")?;
let src_height = NonZeroU32::new(image.height())
.with_context(|| "Failed to get height of the source image")?;
let src_width = image.width();
let src_height = image.height();
let color_type = image.color();
let (src_buffer, pixel_type, mut internal_pixel_type) = match color_type {
ColorType::L8 => (
image.to_luma8().into_raw(),
@@ -189,19 +177,18 @@ fn open_source_image(cli: &Cli) -> Result<(fr::Image<'static>, ColorType, fr::Pi
internal_pixel_type = pixel_type;
}
let mut src_image = fr::Image::from_vec_u8(src_width, src_height, src_buffer, pixel_type)
let mut src_image = Image::from_vec_u8(src_width, src_height, src_buffer, pixel_type)
.with_context(|| "Failed to create source image pixels container")?;
src_image = match cli.colorspace {
structs::ColorSpace::NonLinear => {
debug!("Convert the source image from non-linear colorspace into linear");
let mut linear_src_image =
fr::Image::new(src_image.width(), src_image.height(), internal_pixel_type);
Image::new(src_image.width(), src_image.height(), internal_pixel_type);
if color_type.has_color() {
SRGB_TO_RGB.forward_map(&src_image.view(), &mut linear_src_image.view_mut())?;
SRGB_TO_RGB.forward_map(&src_image, &mut linear_src_image)?;
} else {
GAMMA22_TO_LINEAR
.forward_map(&src_image.view(), &mut linear_src_image.view_mut())?;
GAMMA22_TO_LINEAR.forward_map(&src_image, &mut linear_src_image)?;
}
linear_src_image
}
@@ -209,11 +196,8 @@ fn open_source_image(cli: &Cli) -> Result<(fr::Image<'static>, ColorType, fr::Pi
if internal_pixel_type != pixel_type {
// Convert components of source image into version with high precision
let mut hi_src_image =
fr::Image::new(src_image.width(), src_image.height(), internal_pixel_type);
fr::change_type_of_pixel_components_dyn(
&src_image.view(),
&mut hi_src_image.view_mut(),
)?;
Image::new(src_image.width(), src_image.height(), internal_pixel_type);
fr::change_type_of_pixel_components(&src_image, &mut hi_src_image)?;
hi_src_image
} else {
src_image
@@ -224,24 +208,22 @@ fn open_source_image(cli: &Cli) -> Result<(fr::Image<'static>, ColorType, fr::Pi
Ok((src_image, color_type, pixel_type))
}
fn create_destination_image(cli: &Cli, src_image: &fr::Image) -> fr::Image<'static> {
let aspect_ratio = src_image.width().get() as f32 / src_image.height().get() as f32;
fn create_destination_image(cli: &Cli, src_image: &Image) -> Image<'static> {
if src_image.width() == 0 || src_image.height() == 0 {
return Image::new(0, 0, src_image.pixel_type());
}
let aspect_ratio = src_image.width() as f32 / src_image.height() as f32;
let (dst_width, dst_height) = match (cli.width, cli.height) {
(None, None) => (src_image.width(), src_image.height()),
(Some(width), None) => {
let width = width.calculate_size(src_image.width());
(
width,
get_non_zero_u32((width.get() as f32 / aspect_ratio).round() as u32),
)
(width, (width as f32 / aspect_ratio).round() as u32)
}
(None, Some(height)) => {
let height = height.calculate_size(src_image.height());
(
get_non_zero_u32((height.get() as f32 * aspect_ratio).round() as u32),
height,
)
((height as f32 * aspect_ratio).round() as u32, height)
}
(Some(width), Some(height)) => (
width.calculate_size(src_image.width()),
@@ -249,11 +231,7 @@ fn create_destination_image(cli: &Cli, src_image: &fr::Image) -> fr::Image<'stat
),
};
fr::Image::new(dst_width, dst_height, src_image.pixel_type())
}
fn get_non_zero_u32(v: u32) -> NonZeroU32 {
NonZeroU32::new(v).unwrap_or(NonZeroU32::new(1).unwrap())
Image::new(dst_width, dst_height, src_image.pixel_type())
}
fn get_resizing_algorithm(cli: &Cli) -> fr::ResizeAlg {
@@ -267,7 +245,7 @@ fn get_resizing_algorithm(cli: &Cli) -> fr::ResizeAlg {
fn save_result(
cli: &Cli,
mut image: fr::Image,
mut image: Image,
color_type: ColorType,
pixel_type: fr::PixelType,
) -> Result<()> {
@@ -293,23 +271,18 @@ fn save_result(
image = match cli.colorspace {
structs::ColorSpace::NonLinear => {
debug!("Convert the result image from linear colorspace into non-linear");
let mut non_linear_dst_image =
fr::Image::new(image.width(), image.height(), pixel_type);
let mut non_linear_dst_image = Image::new(image.width(), image.height(), pixel_type);
if color_type.has_color() {
SRGB_TO_RGB.backward_map(&image.view(), &mut non_linear_dst_image.view_mut())?;
SRGB_TO_RGB.backward_map(&image, &mut non_linear_dst_image)?;
} else {
GAMMA22_TO_LINEAR
.backward_map(&image.view(), &mut non_linear_dst_image.view_mut())?;
GAMMA22_TO_LINEAR.backward_map(&image, &mut non_linear_dst_image)?;
}
non_linear_dst_image
}
_ => {
if image.pixel_type() != pixel_type {
let mut lo_src_image = fr::Image::new(image.width(), image.height(), pixel_type);
fr::change_type_of_pixel_components_dyn(
&image.view(),
&mut lo_src_image.view_mut(),
)?;
let mut lo_src_image = Image::new(image.width(), image.height(), pixel_type);
fr::change_type_of_pixel_components(&image, &mut lo_src_image)?;
lo_src_image
} else {
image
@@ -321,8 +294,8 @@ fn save_result(
image::save_buffer(
result_path,
image.buffer(),
image.width().get(),
image.height().get(),
image.width(),
image.height(),
color_type,
)
.with_context(|| "Failed to save the result image")?;
+9 -11
View File
@@ -1,21 +1,19 @@
use crate::get_non_zero_u32;
use fast_image_resize as fr;
use std::num::{NonZeroU16, NonZeroU32, ParseIntError};
use std::num::ParseIntError;
use std::str::FromStr;
use fast_image_resize as fr;
#[derive(Copy, Clone, Debug)]
pub enum Size {
Pixels(NonZeroU32),
Percent(NonZeroU16),
Pixels(u32),
Percent(u16),
}
impl Size {
pub fn calculate_size(&self, src_size: NonZeroU32) -> NonZeroU32 {
pub fn calculate_size(&self, src_size: u32) -> u32 {
match *self {
Self::Pixels(size) => size,
Self::Percent(percent) => get_non_zero_u32(
(src_size.get() as f32 * percent.get() as f32 / 100.).round() as u32,
),
Self::Percent(percent) => (src_size as f32 * percent as f32 / 100.).round() as u32,
}
}
}
@@ -25,9 +23,9 @@ impl FromStr for Size {
fn from_str(s: &str) -> Result<Self, Self::Err> {
if let Some(percent_str) = s.strip_suffix('%') {
NonZeroU16::from_str(percent_str).map(Self::Percent)
u16::from_str(percent_str).map(Self::Percent)
} else {
NonZeroU32::from_str(s).map(Self::Pixels)
u32::from_str(s).map(Self::Pixels)
}
}
}
+5 -10
View File
@@ -1,19 +1,14 @@
use thiserror::Error;
use crate::ImageError;
#[derive(Error, Debug, Clone, Copy)]
#[non_exhaustive]
pub enum MulDivImagesError {
#[error("Source or destination image is not supported")]
ImageError(#[from] ImageError),
#[error("Size of source image does not match to destination image")]
SizeIsDifferent,
#[error("Pixel type of source image does not match to destination image")]
PixelTypeIsDifferent,
#[error("Pixel type of image is not supported")]
UnsupportedPixelType,
}
#[derive(Error, Debug, Clone, Copy)]
#[non_exhaustive]
pub enum MulDivImageError {
#[error("Pixel type of image is not supported")]
UnsupportedPixelType,
PixelTypesAreDifferent,
}
+35 -15
View File
@@ -1,6 +1,4 @@
use crate::pixels::PixelExt;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{pixels, CpuExtensions, ImageError, ImageView, ImageViewMut};
mod common;
pub(crate) mod errors;
@@ -13,29 +11,51 @@ cfg_if::cfg_if! {
}
}
pub(crate) trait AlphaMulDiv
where
Self: PixelExt,
{
pub(crate) trait AlphaMulDiv: pixels::InnerPixel {
/// Multiplies RGB-channels of source image by alpha-channel and store
/// result into destination image.
#[allow(unused_variables)]
fn multiply_alpha(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
);
) -> Result<(), ImageError> {
Err(ImageError::UnsupportedPixelType)
}
/// Multiplies RGB-channels of image by alpha-channel inplace.
fn multiply_alpha_inplace(image: &mut ImageViewMut<Self>, cpu_extensions: CpuExtensions);
#[allow(unused_variables)]
fn multiply_alpha_inplace(
image_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) -> Result<(), ImageError> {
Err(ImageError::UnsupportedPixelType)
}
/// Divides RGB-channels of source image by alpha-channel and store
/// result into destination image.
#[allow(unused_variables)]
fn divide_alpha(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
);
) -> Result<(), ImageError> {
Err(ImageError::UnsupportedPixelType)
}
/// Divides RGB-channels of image by alpha-channel inplace.
fn divide_alpha_inplace(image: &mut ImageViewMut<Self>, cpu_extensions: CpuExtensions);
#[allow(unused_variables)]
fn divide_alpha_inplace(
image_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) -> Result<(), ImageError> {
Err(ImageError::UnsupportedPixelType)
}
}
impl AlphaMulDiv for pixels::U8 {}
impl AlphaMulDiv for pixels::U8x3 {}
impl AlphaMulDiv for pixels::U16 {}
impl AlphaMulDiv for pixels::U16x3 {}
impl AlphaMulDiv for pixels::I32 {}
impl AlphaMulDiv for pixels::F32 {}
+12 -12
View File
@@ -8,11 +8,11 @@ use super::sse4;
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn multiply_alpha(
src_image: &ImageView<U16x2>,
dst_image: &mut ImageViewMut<U16x2>,
src_view: &impl ImageView<Pixel = U16x2>,
dst_view: &mut impl ImageViewMut<Pixel = U16x2>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
@@ -20,8 +20,8 @@ pub(crate) unsafe fn multiply_alpha(
}
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn multiply_alpha_inplace(image: &mut ImageViewMut<U16x2>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x2>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -115,11 +115,11 @@ unsafe fn multiply_alpha_8_pixels(pixels: __m256i) -> __m256i {
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn divide_alpha(
src_image: &ImageView<U16x2>,
dst_image: &mut ImageViewMut<U16x2>,
src_view: &impl ImageView<Pixel = U16x2>,
dst_view: &mut impl ImageViewMut<Pixel = U16x2>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -127,8 +127,8 @@ pub(crate) unsafe fn divide_alpha(
}
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn divide_alpha_inplace(image: &mut ImageViewMut<U16x2>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x2>) {
for row in image_view.iter_rows_mut(0) {
divide_alpha_row_inplace(row);
}
}
+39 -30
View File
@@ -1,6 +1,5 @@
use crate::pixels::U16x2;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageError, ImageView, ImageViewMut};
use super::AlphaMulDiv;
@@ -16,66 +15,76 @@ mod wasm32;
impl AlphaMulDiv for U16x2 {
fn multiply_alpha(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) {
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha(src_image, dst_image) },
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha(src_image, dst_image) },
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::multiply_alpha(src_image, dst_image) },
CpuExtensions::Neon => unsafe { neon::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha(src_image, dst_image) },
_ => native::multiply_alpha(src_image, dst_image),
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha(src_view, dst_view) },
_ => native::multiply_alpha(src_view, dst_view),
}
Ok(())
}
fn multiply_alpha_inplace(image: &mut ImageViewMut<Self>, cpu_extensions: CpuExtensions) {
fn multiply_alpha_inplace(
image_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha_inplace(image) },
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha_inplace(image) },
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::multiply_alpha_inplace(image) },
CpuExtensions::Neon => unsafe { neon::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha_inplace(image) },
_ => native::multiply_alpha_inplace(image),
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha_inplace(image_view) },
_ => native::multiply_alpha_inplace(image_view),
}
Ok(())
}
fn divide_alpha(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) {
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha(src_image, dst_image) },
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha(src_image, dst_image) },
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::divide_alpha(src_image, dst_image) },
CpuExtensions::Neon => unsafe { neon::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha(src_image, dst_image) },
_ => native::divide_alpha(src_image, dst_image),
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha(src_view, dst_view) },
_ => native::divide_alpha(src_view, dst_view),
}
Ok(())
}
fn divide_alpha_inplace(image: &mut ImageViewMut<Self>, cpu_extensions: CpuExtensions) {
fn divide_alpha_inplace(
image_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha_inplace(image) },
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha_inplace(image) },
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::divide_alpha_inplace(image) },
CpuExtensions::Neon => unsafe { neon::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha_inplace(image) },
_ => native::divide_alpha_inplace(image),
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha_inplace(image_view) },
_ => native::divide_alpha_inplace(image_view),
}
Ok(())
}
}
+16 -10
View File
@@ -2,17 +2,20 @@ use crate::alpha::common::{div_and_clip16, mul_div_65535, RECIP_ALPHA16};
use crate::pixels::U16x2;
use crate::{ImageView, ImageViewMut};
pub(crate) fn multiply_alpha(src_image: &ImageView<U16x2>, dst_image: &mut ImageViewMut<U16x2>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) fn multiply_alpha(
src_view: &impl ImageView<Pixel = U16x2>,
dst_view: &mut impl ImageViewMut<Pixel = U16x2>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
}
}
pub(crate) fn multiply_alpha_inplace(image: &mut ImageViewMut<U16x2>) {
for row in image.iter_rows_mut() {
pub(crate) fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x2>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -38,9 +41,12 @@ pub(crate) fn multiply_alpha_row_inplace(row: &mut [U16x2]) {
// Divide
#[inline]
pub(crate) fn divide_alpha(src_image: &ImageView<U16x2>, dst_image: &mut ImageViewMut<U16x2>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) fn divide_alpha(
src_view: &impl ImageView<Pixel = U16x2>,
dst_view: &mut impl ImageViewMut<Pixel = U16x2>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -48,8 +54,8 @@ pub(crate) fn divide_alpha(src_image: &ImageView<U16x2>, dst_image: &mut ImageVi
}
#[inline]
pub(crate) fn divide_alpha_inplace(image: &mut ImageViewMut<U16x2>) {
for row in image.iter_rows_mut() {
pub(crate) fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x2>) {
for row in image_view.iter_rows_mut(0) {
divide_alpha_row_inplace(row);
}
}
+12 -12
View File
@@ -8,11 +8,11 @@ use super::native;
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn multiply_alpha(
src_image: &ImageView<U16x2>,
dst_image: &mut ImageViewMut<U16x2>,
src_view: &impl ImageView<Pixel = U16x2>,
dst_view: &mut impl ImageViewMut<Pixel = U16x2>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
@@ -20,8 +20,8 @@ pub(crate) unsafe fn multiply_alpha(
}
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn multiply_alpha_inplace(image: &mut ImageViewMut<U16x2>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x2>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -111,11 +111,11 @@ unsafe fn multiplies_alpha_4_pixels(pixels: __m128i) -> __m128i {
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn divide_alpha(
src_image: &ImageView<U16x2>,
dst_image: &mut ImageViewMut<U16x2>,
src_view: &impl ImageView<Pixel = U16x2>,
dst_view: &mut impl ImageViewMut<Pixel = U16x2>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -123,8 +123,8 @@ pub(crate) unsafe fn divide_alpha(
}
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn divide_alpha_inplace(image: &mut ImageViewMut<U16x2>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x2>) {
for row in image_view.iter_rows_mut(0) {
divide_alpha_row_inplace(row);
}
}
+12 -12
View File
@@ -8,11 +8,11 @@ use super::sse4;
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn multiply_alpha(
src_image: &ImageView<U16x4>,
dst_image: &mut ImageViewMut<U16x4>,
src_view: &impl ImageView<Pixel = U16x4>,
dst_view: &mut impl ImageViewMut<Pixel = U16x4>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
@@ -20,8 +20,8 @@ pub(crate) unsafe fn multiply_alpha(
}
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn multiply_alpha_inplace(image: &mut ImageViewMut<U16x4>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x4>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -113,11 +113,11 @@ unsafe fn multiply_alpha_4_pixels(pixels: __m256i) -> __m256i {
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn divide_alpha(
src_image: &ImageView<U16x4>,
dst_image: &mut ImageViewMut<U16x4>,
src_view: &impl ImageView<Pixel = U16x4>,
dst_view: &mut impl ImageViewMut<Pixel = U16x4>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -125,8 +125,8 @@ pub(crate) unsafe fn divide_alpha(
}
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn divide_alpha_inplace(image: &mut ImageViewMut<U16x4>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x4>) {
for row in image_view.iter_rows_mut(0) {
divide_alpha_row_inplace(row);
}
}
+39 -30
View File
@@ -1,6 +1,5 @@
use crate::pixels::U16x4;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageError, ImageView, ImageViewMut};
use super::AlphaMulDiv;
@@ -16,66 +15,76 @@ mod wasm32;
impl AlphaMulDiv for U16x4 {
fn multiply_alpha(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) {
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha(src_image, dst_image) },
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha(src_image, dst_image) },
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::multiply_alpha(src_image, dst_image) },
CpuExtensions::Neon => unsafe { neon::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha(src_image, dst_image) },
_ => native::multiply_alpha(src_image, dst_image),
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha(src_view, dst_view) },
_ => native::multiply_alpha(src_view, dst_view),
}
Ok(())
}
fn multiply_alpha_inplace(image: &mut ImageViewMut<Self>, cpu_extensions: CpuExtensions) {
fn multiply_alpha_inplace(
image_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha_inplace(image) },
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha_inplace(image) },
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::multiply_alpha_inplace(image) },
CpuExtensions::Neon => unsafe { neon::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha_inplace(image) },
_ => native::multiply_alpha_inplace(image),
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha_inplace(image_view) },
_ => native::multiply_alpha_inplace(image_view),
}
Ok(())
}
fn divide_alpha(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) {
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha(src_image, dst_image) },
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha(src_image, dst_image) },
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::divide_alpha(src_image, dst_image) },
CpuExtensions::Neon => unsafe { neon::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha(src_image, dst_image) },
_ => native::divide_alpha(src_image, dst_image),
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha(src_view, dst_view) },
_ => native::divide_alpha(src_view, dst_view),
}
Ok(())
}
fn divide_alpha_inplace(image: &mut ImageViewMut<Self>, cpu_extensions: CpuExtensions) {
fn divide_alpha_inplace(
image_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha_inplace(image) },
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha_inplace(image) },
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::divide_alpha_inplace(image) },
CpuExtensions::Neon => unsafe { neon::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha_inplace(image) },
_ => native::divide_alpha_inplace(image),
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha_inplace(image_view) },
_ => native::divide_alpha_inplace(image_view),
}
Ok(())
}
}
+16 -10
View File
@@ -2,17 +2,20 @@ use crate::alpha::common::{div_and_clip16, mul_div_65535, RECIP_ALPHA16};
use crate::pixels::U16x4;
use crate::{ImageView, ImageViewMut};
pub(crate) fn multiply_alpha(src_image: &ImageView<U16x4>, dst_image: &mut ImageViewMut<U16x4>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) fn multiply_alpha(
src_view: &impl ImageView<Pixel = U16x4>,
dst_view: &mut impl ImageViewMut<Pixel = U16x4>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
}
}
pub(crate) fn multiply_alpha_inplace(image: &mut ImageViewMut<U16x4>) {
for row in image.iter_rows_mut() {
pub(crate) fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x4>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -48,9 +51,12 @@ pub(crate) fn multiply_alpha_row_inplace(row: &mut [U16x4]) {
// Divide
#[inline]
pub(crate) fn divide_alpha(src_image: &ImageView<U16x4>, dst_image: &mut ImageViewMut<U16x4>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) fn divide_alpha(
src_view: &impl ImageView<Pixel = U16x4>,
dst_view: &mut impl ImageViewMut<Pixel = U16x4>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -58,8 +64,8 @@ pub(crate) fn divide_alpha(src_image: &ImageView<U16x4>, dst_image: &mut ImageVi
}
#[inline]
pub(crate) fn divide_alpha_inplace(image: &mut ImageViewMut<U16x4>) {
for row in image.iter_rows_mut() {
pub(crate) fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x4>) {
for row in image_view.iter_rows_mut(0) {
divide_alpha_row_inplace(row);
}
}
+12 -12
View File
@@ -8,11 +8,11 @@ use super::native;
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn multiply_alpha(
src_image: &ImageView<U16x4>,
dst_image: &mut ImageViewMut<U16x4>,
src_view: &impl ImageView<Pixel = U16x4>,
dst_view: &mut impl ImageViewMut<Pixel = U16x4>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
@@ -20,8 +20,8 @@ pub(crate) unsafe fn multiply_alpha(
}
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn multiply_alpha_inplace(image: &mut ImageViewMut<U16x4>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x4>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -110,11 +110,11 @@ unsafe fn multiply_alpha_2_pixels(pixels: __m128i) -> __m128i {
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn divide_alpha(
src_image: &ImageView<U16x4>,
dst_image: &mut ImageViewMut<U16x4>,
src_view: &impl ImageView<Pixel = U16x4>,
dst_view: &mut impl ImageViewMut<Pixel = U16x4>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -122,8 +122,8 @@ pub(crate) unsafe fn divide_alpha(
}
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn divide_alpha_inplace(image: &mut ImageViewMut<U16x4>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U16x4>) {
for row in image_view.iter_rows_mut(0) {
divide_alpha_row_inplace(row);
}
}
+15 -13
View File
@@ -1,19 +1,18 @@
use std::arch::x86_64::*;
use crate::pixels::U8x2;
use crate::simd_utils;
use crate::utils::foreach_with_pre_reading;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
use super::sse4;
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn multiply_alpha(
src_image: &ImageView<U8x2>,
dst_image: &mut ImageViewMut<U8x2>,
src_view: &impl ImageView<Pixel = U8x2>,
dst_view: &mut impl ImageViewMut<Pixel = U8x2>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
@@ -21,8 +20,8 @@ pub(crate) unsafe fn multiply_alpha(
}
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn multiply_alpha_inplace(image: &mut ImageViewMut<U8x2>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x2>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -112,9 +111,12 @@ unsafe fn multiply_alpha_16_pixels(pixels: __m256i) -> __m256i {
// Divide
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn divide_alpha(src_image: &ImageView<U8x2>, dst_image: &mut ImageViewMut<U8x2>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) unsafe fn divide_alpha(
src_view: &impl ImageView<Pixel = U8x2>,
dst_view: &mut impl ImageViewMut<Pixel = U8x2>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -122,8 +124,8 @@ pub(crate) unsafe fn divide_alpha(src_image: &ImageView<U8x2>, dst_image: &mut I
}
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn divide_alpha_inplace(image: &mut ImageViewMut<U8x2>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x2>) {
for row in image_view.iter_rows_mut(0) {
divide_alpha_row_inplace(row);
}
}
+40 -30
View File
@@ -1,6 +1,6 @@
use crate::cpu_extensions::CpuExtensions;
use crate::pixels::U8x2;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{ImageError, ImageView, ImageViewMut};
use super::AlphaMulDiv;
@@ -16,66 +16,76 @@ mod wasm32;
impl AlphaMulDiv for U8x2 {
fn multiply_alpha(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) {
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha(src_image, dst_image) },
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha(src_image, dst_image) },
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::multiply_alpha(src_image, dst_image) },
CpuExtensions::Neon => unsafe { neon::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha(src_image, dst_image) },
_ => native::multiply_alpha(src_image, dst_image),
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha(src_view, dst_view) },
_ => native::multiply_alpha(src_view, dst_view),
}
Ok(())
}
fn multiply_alpha_inplace(image: &mut ImageViewMut<Self>, cpu_extensions: CpuExtensions) {
fn multiply_alpha_inplace(
image_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha_inplace(image) },
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha_inplace(image) },
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::multiply_alpha_inplace(image) },
CpuExtensions::Neon => unsafe { neon::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha_inplace(image) },
_ => native::multiply_alpha_inplace(image),
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha_inplace(image_view) },
_ => native::multiply_alpha_inplace(image_view),
}
Ok(())
}
fn divide_alpha(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) {
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha(src_image, dst_image) },
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha(src_image, dst_image) },
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::divide_alpha(src_image, dst_image) },
CpuExtensions::Neon => unsafe { neon::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha(src_image, dst_image) },
_ => native::divide_alpha(src_image, dst_image),
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha(src_view, dst_view) },
_ => native::divide_alpha(src_view, dst_view),
}
Ok(())
}
fn divide_alpha_inplace(image: &mut ImageViewMut<Self>, cpu_extensions: CpuExtensions) {
fn divide_alpha_inplace(
image_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha_inplace(image) },
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha_inplace(image) },
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::divide_alpha_inplace(image) },
CpuExtensions::Neon => unsafe { neon::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha_inplace(image) },
_ => native::divide_alpha_inplace(image),
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha_inplace(image_view) },
_ => native::divide_alpha_inplace(image_view),
}
Ok(())
}
}
+22 -16
View File
@@ -2,17 +2,20 @@ use crate::alpha::common::{div_and_clip, mul_div_255, RECIP_ALPHA};
use crate::pixels::U8x2;
use crate::{ImageView, ImageViewMut};
pub(crate) fn multiply_alpha(src_image: &ImageView<U8x2>, dst_image: &mut ImageViewMut<U8x2>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) fn multiply_alpha(
src_view: &impl ImageView<Pixel = U8x2>,
dst_view: &mut impl ImageViewMut<Pixel = U8x2>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
}
}
pub(crate) fn multiply_alpha_inplace(image: &mut ImageViewMut<U8x2>) {
for row in image.iter_rows_mut() {
pub(crate) fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x2>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -20,27 +23,30 @@ pub(crate) fn multiply_alpha_inplace(image: &mut ImageViewMut<U8x2>) {
#[inline(always)]
pub(crate) fn multiply_alpha_row(src_row: &[U8x2], dst_row: &mut [U8x2]) {
for (src_pixel, dst_pixel) in src_row.iter().zip(dst_row) {
let components: [u8; 2] = src_pixel.0.to_le_bytes();
let components: [u8; 2] = src_pixel.0;
let alpha = components[1];
dst_pixel.0 = u16::from_le_bytes([mul_div_255(components[0], alpha), alpha]);
dst_pixel.0 = [mul_div_255(components[0], alpha), alpha];
}
}
#[inline(always)]
pub(crate) fn multiply_alpha_row_inplace(row: &mut [U8x2]) {
for pixel in row {
let components: [u8; 2] = pixel.0.to_le_bytes();
let components: [u8; 2] = pixel.0;
let alpha = components[1];
pixel.0 = u16::from_le_bytes([mul_div_255(components[0], alpha), alpha]);
pixel.0 = [mul_div_255(components[0], alpha), alpha];
}
}
// Divide
#[inline]
pub(crate) fn divide_alpha(src_image: &ImageView<U8x2>, dst_image: &mut ImageViewMut<U8x2>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) fn divide_alpha(
src_view: &impl ImageView<Pixel = U8x2>,
dst_view: &mut impl ImageViewMut<Pixel = U8x2>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -48,8 +54,8 @@ pub(crate) fn divide_alpha(src_image: &ImageView<U8x2>, dst_image: &mut ImageVie
}
#[inline]
pub(crate) fn divide_alpha_inplace(image: &mut ImageViewMut<U8x2>) {
for dst_row in image.iter_rows_mut() {
pub(crate) fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x2>) {
for dst_row in image_view.iter_rows_mut(0) {
let src_row = unsafe { std::slice::from_raw_parts(dst_row.as_ptr(), dst_row.len()) };
divide_alpha_row(src_row, dst_row);
}
@@ -61,9 +67,9 @@ pub(crate) fn divide_alpha_row(src_row: &[U8x2], dst_row: &mut [U8x2]) {
.iter()
.zip(dst_row)
.for_each(|(src_pixel, dst_pixel)| {
let components: [u8; 2] = src_pixel.0.to_le_bytes();
let components: [u8; 2] = src_pixel.0;
let alpha = components[1];
let recip_alpha = RECIP_ALPHA[alpha as usize];
dst_pixel.0 = u16::from_le_bytes([div_and_clip(components[0], recip_alpha), alpha]);
dst_pixel.0 = [div_and_clip(components[0], recip_alpha), alpha];
});
}
+18 -15
View File
@@ -8,11 +8,11 @@ use super::native;
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn multiply_alpha(
src_image: &ImageView<U8x2>,
dst_image: &mut ImageViewMut<U8x2>,
src_view: &impl ImageView<Pixel = U8x2>,
dst_view: &mut impl ImageViewMut<Pixel = U8x2>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
@@ -20,8 +20,8 @@ pub(crate) unsafe fn multiply_alpha(
}
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn multiply_alpha_inplace(image: &mut ImageViewMut<U8x2>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x2>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -103,9 +103,12 @@ unsafe fn multiplies_alpha_8_pixels(pixels: __m128i) -> __m128i {
// Divide
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn divide_alpha(src_image: &ImageView<U8x2>, dst_image: &mut ImageViewMut<U8x2>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) unsafe fn divide_alpha(
src_view: &impl ImageView<Pixel = U8x2>,
dst_view: &mut impl ImageViewMut<Pixel = U8x2>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -113,8 +116,8 @@ pub(crate) unsafe fn divide_alpha(src_image: &ImageView<U8x2>, dst_image: &mut I
}
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn divide_alpha_inplace(image: &mut ImageViewMut<U8x2>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x2>) {
for row in image_view.iter_rows_mut(0) {
divide_alpha_row_inplace(row);
}
}
@@ -141,13 +144,13 @@ pub(crate) unsafe fn divide_alpha_row(src_row: &[U8x2], dst_row: &mut [U8x2]) {
if !src_remainder.is_empty() {
let dst_reminder = dst_chunks.into_remainder();
let mut src_pixels = [U8x2::new(0); 8];
let mut src_pixels = [U8x2::new([0; 2]); 8];
src_pixels
.iter_mut()
.zip(src_remainder)
.for_each(|(d, s)| *d = *s);
let mut dst_pixels = [U8x2::new(0); 8];
let mut dst_pixels = [U8x2::new([0; 2]); 8];
let mut pixels = _mm_loadu_si128(src_pixels.as_ptr() as *const __m128i);
pixels = divide_alpha_8_pixels(pixels);
_mm_storeu_si128(dst_pixels.as_mut_ptr() as *mut __m128i, pixels);
@@ -178,13 +181,13 @@ pub(crate) unsafe fn divide_alpha_row_inplace(row: &mut [U8x2]) {
let reminder = chunks.into_remainder();
if !reminder.is_empty() {
let mut src_pixels = [U8x2::new(0); 8];
let mut src_pixels = [U8x2::new([0; 2]); 8];
src_pixels
.iter_mut()
.zip(reminder.iter())
.for_each(|(d, s)| *d = *s);
let mut dst_pixels = [U8x2::new(0); 8];
let mut dst_pixels = [U8x2::new([0; 2]); 8];
let mut pixels = _mm_loadu_si128(src_pixels.as_ptr() as *const __m128i);
pixels = divide_alpha_8_pixels(pixels);
_mm_storeu_si128(dst_pixels.as_mut_ptr() as *mut __m128i, pixels);
+15 -13
View File
@@ -1,27 +1,26 @@
use std::arch::x86_64::*;
use crate::pixels::U8x4;
use crate::simd_utils;
use crate::utils::foreach_with_pre_reading;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
use super::sse4;
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn multiply_alpha(
src_image: &ImageView<U8x4>,
dst_image: &mut ImageViewMut<U8x4>,
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
}
}
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn multiply_alpha_inplace(image: &mut ImageViewMut<U8x4>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x4>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -108,9 +107,12 @@ unsafe fn multiply_alpha_8_pixels(pixels: __m256i) -> __m256i {
// Divide
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn divide_alpha(src_image: &ImageView<U8x4>, dst_image: &mut ImageViewMut<U8x4>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) unsafe fn divide_alpha(
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -118,8 +120,8 @@ pub(crate) unsafe fn divide_alpha(src_image: &ImageView<U8x4>, dst_image: &mut I
}
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn divide_alpha_inplace(image: &mut ImageViewMut<U8x4>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x4>) {
for row in image_view.iter_rows_mut(0) {
divide_alpha_row_inplace(row);
}
}
+35 -26
View File
@@ -1,6 +1,5 @@
use crate::pixels::U8x4;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageError, ImageView, ImageViewMut};
use super::AlphaMulDiv;
@@ -16,66 +15,76 @@ mod wasm32;
impl AlphaMulDiv for U8x4 {
fn multiply_alpha(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) {
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha(src_image, dst_image) },
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha(src_image, dst_image) },
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha(src_view, dst_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::multiply_alpha(src_image, dst_image) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha(src_image, dst_image) },
_ => native::multiply_alpha(src_image, dst_image),
_ => native::multiply_alpha(src_view, dst_view),
}
Ok(())
}
fn multiply_alpha_inplace(image: &mut ImageViewMut<Self>, cpu_extensions: CpuExtensions) {
fn multiply_alpha_inplace(
image_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha_inplace(image) },
CpuExtensions::Avx2 => unsafe { avx2::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha_inplace(image) },
CpuExtensions::Sse4_1 => unsafe { sse4::multiply_alpha_inplace(image_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::multiply_alpha_inplace(image) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::multiply_alpha_inplace(image) },
_ => native::multiply_alpha_inplace(image),
_ => native::multiply_alpha_inplace(image_view),
}
Ok(())
}
fn divide_alpha(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) {
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha(src_image, dst_image) },
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha(src_image, dst_image) },
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::divide_alpha(src_image, dst_image) },
CpuExtensions::Neon => unsafe { neon::divide_alpha(src_view, dst_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha(src_image, dst_image) },
_ => native::divide_alpha(src_image, dst_image),
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha(src_view, dst_view) },
_ => native::divide_alpha(src_view, dst_view),
}
Ok(())
}
fn divide_alpha_inplace(image: &mut ImageViewMut<Self>, cpu_extensions: CpuExtensions) {
fn divide_alpha_inplace(
image_view: &mut impl ImageViewMut<Pixel = Self>,
cpu_extensions: CpuExtensions,
) -> Result<(), ImageError> {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha_inplace(image) },
CpuExtensions::Avx2 => unsafe { avx2::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha_inplace(image) },
CpuExtensions::Sse4_1 => unsafe { sse4::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => unsafe { neon::divide_alpha_inplace(image) },
CpuExtensions::Neon => unsafe { neon::divide_alpha_inplace(image_view) },
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha_inplace(image) },
_ => native::divide_alpha_inplace(image),
CpuExtensions::Simd128 => unsafe { wasm32::divide_alpha_inplace(image_view) },
_ => native::divide_alpha_inplace(image_view),
}
Ok(())
}
}
+16 -10
View File
@@ -2,9 +2,12 @@ use crate::alpha::common::{div_and_clip, mul_div_255, RECIP_ALPHA};
use crate::pixels::U8x4;
use crate::{ImageView, ImageViewMut};
pub(crate) fn multiply_alpha(src_image: &ImageView<U8x4>, dst_image: &mut ImageViewMut<U8x4>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) fn multiply_alpha(
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
for (src_pixel, dst_pixel) in src_row.iter().zip(dst_row.iter_mut()) {
@@ -13,8 +16,8 @@ pub(crate) fn multiply_alpha(src_image: &ImageView<U8x4>, dst_image: &mut ImageV
}
}
pub(crate) fn multiply_alpha_inplace(image: &mut ImageViewMut<U8x4>) {
for row in image.iter_rows_mut() {
pub(crate) fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x4>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -48,9 +51,12 @@ fn multiply_alpha_pixel(mut pixel: U8x4) -> U8x4 {
// Divide
#[inline]
pub(crate) fn divide_alpha(src_image: &ImageView<U8x4>, dst_image: &mut ImageViewMut<U8x4>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) fn divide_alpha(
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
@@ -58,8 +64,8 @@ pub(crate) fn divide_alpha(src_image: &ImageView<U8x4>, dst_image: &mut ImageVie
}
#[inline]
pub(crate) fn divide_alpha_inplace(image: &mut ImageViewMut<U8x4>) {
for row in image.iter_rows_mut() {
pub(crate) fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x4>) {
for row in image_view.iter_rows_mut(0) {
row.iter_mut().for_each(|pixel| {
*pixel = divide_alpha_pixel(*pixel);
});
+14 -11
View File
@@ -8,11 +8,11 @@ use super::native;
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn multiply_alpha(
src_image: &ImageView<U8x4>,
dst_image: &mut ImageViewMut<U8x4>,
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
multiply_alpha_row(src_row, dst_row);
@@ -20,8 +20,8 @@ pub(crate) unsafe fn multiply_alpha(
}
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn multiply_alpha_inplace(image: &mut ImageViewMut<U8x4>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn multiply_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x4>) {
for row in image_view.iter_rows_mut(0) {
multiply_alpha_row_inplace(row);
}
}
@@ -100,17 +100,20 @@ unsafe fn multiply_alpha_4_pixels(pixels: __m128i) -> __m128i {
// Divide
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn divide_alpha(src_image: &ImageView<U8x4>, dst_image: &mut ImageViewMut<U8x4>) {
let src_rows = src_image.iter_rows(0);
let dst_rows = dst_image.iter_rows_mut();
pub(crate) unsafe fn divide_alpha(
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
) {
let src_rows = src_view.iter_rows(0);
let dst_rows = dst_view.iter_rows_mut(0);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
divide_alpha_row(src_row, dst_row);
}
}
#[target_feature(enable = "sse4.1")]
pub(crate) unsafe fn divide_alpha_inplace(image: &mut ImageViewMut<U8x4>) {
for row in image.iter_rows_mut() {
pub(crate) unsafe fn divide_alpha_inplace(image_view: &mut impl ImageViewMut<Pixel = U8x4>) {
for row in image_view.iter_rows_mut(0) {
divide_alpha_row_inplace(row);
}
}
+234
View File
@@ -0,0 +1,234 @@
use core::array::IntoIter;
use core::iter::Take;
use core::iter::{FusedIterator, Iterator};
use std::mem::MaybeUninit;
/// An iterator over `N` elements of the iterator at a time.
///
/// The chunks do not overlap. If `N` does not divide the length of the
/// iterator, then the last up to `N-1` elements will be omitted.
#[derive(Debug, Clone)]
#[must_use = "iterators are lazy and do nothing unless consumed"]
pub struct ArrayChunks<I: Iterator, const N: usize> {
iter: I,
remainder: Option<Take<IntoIter<I::Item, N>>>,
}
impl<I, const N: usize> ArrayChunks<I, N>
where
I: Iterator,
{
pub fn new(iter: I) -> Self {
assert_ne!(N, 0, "chunk size must be non-zero");
Self {
iter,
remainder: None,
}
}
/// Returns an iterator over the remaining elements of the original iterator
/// that are not going to be returned by this iterator. The returned
/// iterator will yield at most `N-1` elements.
#[inline]
pub fn into_remainder(self) -> Option<Take<IntoIter<I::Item, N>>> {
self.remainder
}
}
impl<I, const N: usize> Iterator for ArrayChunks<I, N>
where
I: Iterator,
{
type Item = [I::Item; N];
#[inline]
fn next(&mut self) -> Option<Self::Item> {
match next_chunk(&mut self.iter) {
Ok(chunk) => Some(chunk),
Err(remainder) => {
// Make sure to not override `self.remainder` with an empty array
// when `next` is called after `ArrayChunks` exhaustion.
self.remainder.get_or_insert(remainder);
None
}
}
// self.try_for_each(ControlFlow::Break).break_value()
}
#[inline]
fn size_hint(&self) -> (usize, Option<usize>) {
let (lower, upper) = self.iter.size_hint();
(lower / N, upper.map(|n| n / N))
}
#[inline]
fn count(self) -> usize {
self.iter.count() / N
}
}
#[inline]
fn next_chunk<I: Iterator, const N: usize>(
iter: &mut I,
) -> Result<[I::Item; N], Take<IntoIter<I::Item, N>>>
where
I: Sized,
{
iter_next_chunk(iter)
}
impl<I, const N: usize> FusedIterator for ArrayChunks<I, N> where I: FusedIterator {}
impl<I, const N: usize> ExactSizeIterator for ArrayChunks<I, N>
where
I: ExactSizeIterator,
{
#[inline]
fn len(&self) -> usize {
self.iter.len() / N
}
}
/// Pulls `N` items from `iter` and returns them as an array. If the iterator
/// yields fewer than `N` items, `Err` is returned containing an iterator over
/// the already yielded items.
///
/// Since the iterator is passed as a mutable reference and this function calls
/// `next` at most `N` times, the iterator can still be used afterwards to
/// retrieve the remaining items.
///
/// If `iter.next()` panicks, all items already yielded by the iterator are
/// dropped.
///
/// Used for [`Iterator::next_chunk`].
#[inline]
fn iter_next_chunk<T, const N: usize>(
iter: &mut impl Iterator<Item = T>,
) -> Result<[T; N], Take<IntoIter<T, N>>> {
let mut array = uninit_array::<T, N>();
let r = iter_next_chunk_erased(&mut array, iter);
match r {
Ok(()) => {
// SAFETY: All elements of `array` were populated.
Ok(unsafe { array_assume_init(array) })
}
Err(initialized) => {
// SAFETY: Only the first `initialized` elements were populated
let array = unsafe { array_assume_init(array) };
Err(array.into_iter().take(initialized))
// Err(unsafe { IntoIter::new_unchecked(array, 0..initialized) })
}
}
}
#[inline(always)]
const fn uninit_array<T, const N: usize>() -> [MaybeUninit<T>; N] {
// SAFETY: An uninitialized `[MaybeUninit<_>; LEN]` is valid.
unsafe { MaybeUninit::<[MaybeUninit<T>; N]>::uninit().assume_init() }
}
#[inline(always)]
unsafe fn array_assume_init<T, const N: usize>(array: [MaybeUninit<T>; N]) -> [T; N] {
// SAFETY:
// * The caller guarantees that all elements of the array are initialized
// * `MaybeUninit<T>` and T are guaranteed to have the same layout
// * `MaybeUninit` does not drop, so there are no double-frees
// And thus the conversion is safe
let ret = unsafe {
// core::intrinsics::assert_inhabited::<[T; N]>();
(&array as *const _ as *const [T; N]).read()
};
// FIXME: required to avoid `~const Destruct` bound
core::mem::forget(array);
ret
}
/// Version of [`iter_next_chunk`] using a passed-in slice in order to avoid
/// needing to monomorphize for every array length.
///
/// Unfortunately this loop has two exit conditions, the buffer filling up
/// or the iterator running out of items, making it tend to optimize poorly.
#[inline]
fn iter_next_chunk_erased<T>(
buffer: &mut [MaybeUninit<T>],
iter: &mut impl Iterator<Item = T>,
) -> Result<(), usize> {
let mut guard = Guard {
array_mut: buffer,
initialized: 0,
};
while guard.initialized < guard.array_mut.len() {
let Some(item) = iter.next() else {
// Unlike `try_from_fn_erased`, we want to keep the partial results,
// so we need to defuse the guard instead of using `?`.
let initialized = guard.initialized;
core::mem::forget(guard);
return Err(initialized);
};
// SAFETY: The loop condition ensures we have space to push the item
unsafe { guard.push_unchecked(item) };
}
core::mem::forget(guard);
Ok(())
}
/// Panic guard for incremental initialization of arrays.
///
/// Disarm the guard with `mem::forget` once the array has been initialized.
///
/// # Safety
///
/// All write accesses to this structure are unsafe and must maintain a correct
/// count of `initialized` elements.
///
/// To minimize indirection fields are still pub but callers should at least use
/// `push_unchecked` to signal that something unsafe is going on.
struct Guard<'a, T> {
/// The array to be initialized.
pub array_mut: &'a mut [MaybeUninit<T>],
/// The number of items that have been initialized so far.
pub initialized: usize,
}
impl<T> Guard<'_, T> {
/// Adds an item to the array and updates the initialized item counter.
///
/// # Safety
///
/// No more than N elements must be initialized.
#[inline]
pub unsafe fn push_unchecked(&mut self, item: T) {
// SAFETY: If `initialized` was correct before and the caller does not
// invoke this method more than N times then writes will be in-bounds
// and slots will not be initialized more than once.
unsafe {
self.array_mut
.get_unchecked_mut(self.initialized)
.write(item);
self.initialized += 1;
}
}
}
impl<T> Drop for Guard<'_, T> {
fn drop(&mut self) {
debug_assert!(self.initialized <= self.array_mut.len());
// SAFETY: this slice will contain only initialized objects.
unsafe {
core::ptr::drop_in_place(slice_assume_init_mut(
self.array_mut.get_unchecked_mut(..self.initialized),
));
}
}
}
#[inline(always)]
unsafe fn slice_assume_init_mut<T>(slice: &mut [MaybeUninit<T>]) -> &mut [T] {
// SAFETY: similar to safety notes for `slice_get_ref`, but we have a
// mutable reference which is also guaranteed to be valid for writes.
unsafe { &mut *(slice as *mut [MaybeUninit<T>] as *mut [T]) }
}
+84
View File
@@ -0,0 +1,84 @@
use crate::pixels::{
InnerPixel, IntoPixelComponent, U16x2, U16x3, U16x4, U8x2, U8x3, U8x4, U16, U8,
};
use crate::{
try_pixel_type, DifferentDimensionsError, ImageView, ImageViewMut, IntoImageView,
IntoImageViewMut, MappingError, PixelType,
};
pub fn change_type_of_pixel_components(
src_image: &impl IntoImageView,
dst_image: &mut impl IntoImageViewMut,
) -> Result<(), MappingError> {
macro_rules! map {
($value:expr, $(($low_enum:path, $low_pt:ty, $high_enum:path, $high_pt:ty)),*) => {
match $value {
$(
($low_enum, $low_enum) =>
change_components_type::<$low_pt, $low_pt>(src_image, dst_image),
($low_enum, $high_enum) =>
change_components_type::<$low_pt, $high_pt>(src_image, dst_image),
($high_enum, $low_enum) =>
change_components_type::<$high_pt, $low_pt>(src_image, dst_image),
($high_enum, $high_enum) =>
change_components_type::<$high_pt, $high_pt>(src_image, dst_image),
)*
_ => Err(MappingError::UnsupportedCombinationOfImageTypes),
}
}
}
let src_pixel_type = try_pixel_type(src_image)?;
let dst_pixel_type = try_pixel_type(dst_image)?;
use PixelType as PT;
map!(
(src_pixel_type, dst_pixel_type),
(PT::U8, U8, PT::U16, U16),
(PT::U8x2, U8x2, PT::U16x2, U16x2),
(PT::U8x3, U8x3, PT::U16x3, U16x3),
(PT::U8x4, U8x4, PT::U16x4, U16x4)
)
}
#[inline(always)]
fn change_components_type<S, D>(
src_image: &impl IntoImageView,
dst_image: &mut impl IntoImageViewMut,
) -> Result<(), MappingError>
where
S: InnerPixel,
D: InnerPixel<CountOfComponents = S::CountOfComponents>,
<S as InnerPixel>::Component: IntoPixelComponent<<D as InnerPixel>::Component>,
{
match (src_image.image_view::<S>(), dst_image.image_view_mut::<D>()) {
(Some(src_view), Some(mut dst_view)) => {
change_type_of_pixel_components_typed(&src_view, &mut dst_view).map_err(|e| e.into())
}
_ => Err(MappingError::UnsupportedCombinationOfImageTypes),
}
}
pub fn change_type_of_pixel_components_typed<S, D>(
src_image: &impl ImageView<Pixel = S>,
dst_image: &mut impl ImageViewMut<Pixel = D>,
) -> Result<(), DifferentDimensionsError>
where
S: InnerPixel,
D: InnerPixel<CountOfComponents = S::CountOfComponents>,
<S as InnerPixel>::Component: IntoPixelComponent<<D as InnerPixel>::Component>,
{
if src_image.width() != dst_image.width() || src_image.height() != dst_image.height() {
return Err(DifferentDimensionsError);
}
for (s_row, d_row) in src_image.iter_rows(0).zip(dst_image.iter_rows_mut(0)) {
let s_components = S::components(s_row);
let d_components = D::components_mut(d_row);
for (&s_comp, d_comp) in s_components.iter().zip(d_components) {
*d_comp = s_comp.into_component();
}
}
Ok(())
}
+114 -55
View File
@@ -2,9 +2,13 @@
use num_traits::bounds::UpperBounded;
use num_traits::Zero;
use crate::pixels::{GetCount, IntoPixelComponent, PixelComponent, PixelExt, Values};
use crate::{DynamicImageView, DynamicImageViewMut, MappingError};
use crate::{ImageView, ImageViewMut};
use crate::pixels::{
GetCount, InnerPixel, IntoPixelComponent, PixelComponent, PixelType, U16x2, U16x3, U16x4, U8x2,
U8x3, U8x4, Values, U16, U8,
};
use crate::{
try_pixel_type, ImageView, ImageViewMut, IntoImageView, IntoImageViewMut, MappingError,
};
pub(crate) mod mappers;
@@ -86,19 +90,43 @@ where
}
}
pub fn map_image<S, D, In, CC>(&self, src_image: &ImageView<S>, dst_image: &mut ImageViewMut<D>)
pub fn map_image<S, D>(
&self,
src_image: &impl IntoImageView,
dst_image: &mut impl IntoImageViewMut,
) -> Result<(), MappingError>
where
In: PixelComponent<CountOfComponentValues = Values<SIZE>>
S: InnerPixel,
<S as InnerPixel>::Component: PixelComponent<CountOfComponentValues = Values<SIZE>>
+ IntoPixelComponent<Out>
+ Into<usize>,
CC: GetCount,
S: PixelExt<Component = In, CountOfComponents = CC>,
D: PixelExt<Component = Out, CountOfComponents = CC>,
D: InnerPixel<Component = Out, CountOfComponents = S::CountOfComponents>,
{
for (s_row, d_row) in src_image.iter_rows(0).zip(dst_image.iter_rows_mut()) {
let (src_view, dst_view) =
match (src_image.image_view::<S>(), dst_image.image_view_mut::<D>()) {
(Some(src_view), Some(dst_view)) => (src_view, dst_view),
_ => return Err(MappingError::UnsupportedCombinationOfImageTypes),
};
self.map_image_typed(src_view, dst_view);
Ok(())
}
pub fn map_image_typed<S, D>(
&self,
src_view: impl ImageView<Pixel = S>,
mut dst_view: impl ImageViewMut<Pixel = D>,
) where
S: InnerPixel,
<S as InnerPixel>::Component: PixelComponent<CountOfComponentValues = Values<SIZE>>
+ IntoPixelComponent<Out>
+ Into<usize>,
D: InnerPixel<Component = Out, CountOfComponents = S::CountOfComponents>,
{
for (s_row, d_row) in src_view.iter_rows(0).zip(dst_view.iter_rows_mut(0)) {
let s_comp = S::components(s_row);
let d_comp = D::components_mut(d_row);
match CC::count() {
match S::CountOfComponents::count() {
2 => self.map_with_gaps(s_comp, d_comp, 2), // Don't map alpha channel
4 => self.map_with_gaps(s_comp, d_comp, 4), // Don't map alpha channel
_ => self.map(s_comp, d_comp),
@@ -106,15 +134,30 @@ where
}
}
pub fn map_image_inplace<S, CC>(&self, image: &mut ImageViewMut<S>)
pub fn map_image_inplace<S>(
&self,
image: &mut impl IntoImageViewMut,
) -> Result<(), MappingError>
where
CC: GetCount,
S: PixelExt<Component = Out, CountOfComponents = CC>,
Out: Into<usize>,
S: InnerPixel<Component = Out>,
{
for row in image.iter_rows_mut() {
if let Some(image_view) = image.image_view_mut::<S>() {
self.map_image_inplace_typed(image_view);
Ok(())
} else {
Err(MappingError::UnsupportedCombinationOfImageTypes)
}
}
pub fn map_image_inplace_typed<S>(&self, mut image_view: impl ImageViewMut<Pixel = S>)
where
Out: Into<usize>,
S: InnerPixel<Component = Out>,
{
for row in image_view.iter_rows_mut(0) {
let comp = S::components_mut(row);
match CC::count() {
match S::CountOfComponents::count() {
2 => self.map_with_gaps_inplace(comp, 2), // Don't map alpha channel
4 => self.map_with_gaps_inplace(comp, 4), // Don't map alpha channel
_ => self.map_inplace(comp),
@@ -135,11 +178,12 @@ struct MappingTablesGroup {
/// This structure holds tables for mapping values of pixel's
/// components in forward and backward directions.
///
/// Supported all pixel types exclude `I32` and `F32`.
/// All pixel types except `I32` and `F32` are supported.
///
/// Source and destination images may have different bit depth of one pixel component.
/// Source and destination images may have different bit depth of one
/// pixel component.
/// But count of components must be equal.
/// For example, you may convert `U8x3` image with sRGB colorspace into
/// For example, you can convert `U8x3` image with sRGB colorspace into
/// `U16x3` image with linear colorspace.
///
/// Alpha channel from such pixel types as `U8x2`, `U8x4`, `U16x2` and `U16x4`
@@ -198,27 +242,41 @@ impl PixelComponentMapper {
fn map(
tables: &MappingTablesGroup,
src_image: &DynamicImageView,
dst_image: &mut DynamicImageViewMut,
src_image: &impl IntoImageView,
dst_image: &mut impl IntoImageViewMut,
) -> Result<(), MappingError> {
let src_pixel_type = try_pixel_type(src_image)?;
let dst_pixel_type = try_pixel_type(dst_image)?;
if src_image.width() != dst_image.width() || src_image.height() != dst_image.height() {
return Err(MappingError::DifferentDimensions);
}
use DynamicImageView as DI;
use DynamicImageViewMut as DIMut;
use PixelType as PT;
macro_rules! match_img {
(
$tables: ident, $src_image: ident, $dst_image: ident,
$(($p8: path, $p16: path, $p8_mut: path, $p16_mut: path),)*
$tables: ident,
$(($p8: path, $pt8: tt, $p16: path, $pt16: tt),)*
) => {
match ($src_image, $dst_image) {
match (src_pixel_type, dst_pixel_type) {
$(
($p8(src), $p8_mut(dst)) => $tables.u8_u8.map_image(src, dst),
($p8(src), $p16_mut(dst)) => $tables.u8_u16.map_image(src, dst),
($p16(src), $p8_mut(dst)) => $tables.u16_u8.map_image(src, dst),
($p16(src), $p16_mut(dst)) => $tables.u16_u16.map_image(src, dst),
($p8, $p8) => $tables.u8_u8.map_image::<$pt8, $pt8>(
src_image,
dst_image,
),
($p8, $p16) => $tables.u8_u16.map_image::<$pt8, $pt16>(
src_image,
dst_image,
),
($p16, $p8) => $tables.u16_u8.map_image::<$pt16, $pt8>(
src_image,
dst_image,
),
($p16, $p16) => $tables.u16_u16.map_image::<$pt16, $pt16>(
src_image,
dst_image,
),
)*
_ => return Err(MappingError::UnsupportedCombinationOfImageTypes),
}
@@ -227,31 +285,30 @@ impl PixelComponentMapper {
match_img!(
tables,
src_image,
dst_image,
(DI::U8, DI::U16, DIMut::U8, DIMut::U16),
(DI::U8x2, DI::U16x2, DIMut::U8x2, DIMut::U16x2),
(DI::U8x3, DI::U16x3, DIMut::U8x3, DIMut::U16x3),
(DI::U8x4, DI::U16x4, DIMut::U8x4, DIMut::U16x4),
);
Ok(())
(PT::U8, U8, PT::U16, U16),
(PT::U8x2, U8x2, PT::U16x2, U16x2),
(PT::U8x3, U8x3, PT::U16x3, U16x3),
(PT::U8x4, U8x4, PT::U16x4, U16x4),
)
}
fn map_inplace(
tables: &MappingTablesGroup,
image: &mut DynamicImageViewMut,
image: &mut impl IntoImageViewMut,
) -> Result<(), MappingError> {
use DynamicImageViewMut as DIMut;
let pixel_type = try_pixel_type(image)?;
use PixelType as PT;
macro_rules! match_img {
(
$tables: ident, $image: ident,
$(($p8_mut: path, $p16_mut: path),)*
$(($p8: path, $pt8: tt, $p16: path, $pt16: tt),)*
) => {
match $image {
match pixel_type {
$(
$p8_mut(img) => $tables.u8_u8.map_image_inplace(img),
$p16_mut(img) => $tables.u16_u16.map_image_inplace(img),
$p8 => $tables.u8_u8.map_image_inplace::<$pt8>($image),
$p16 => $tables.u16_u16.map_image_inplace::<$pt16>($image),
)*
_ => return Err(MappingError::UnsupportedCombinationOfImageTypes),
}
@@ -261,25 +318,27 @@ impl PixelComponentMapper {
match_img!(
tables,
image,
(DIMut::U8, DIMut::U16),
(DIMut::U8x2, DIMut::U16x2),
(DIMut::U8x3, DIMut::U16x3),
(DIMut::U8x4, DIMut::U16x4),
);
Ok(())
(PT::U8, U8, PT::U16, U16),
(PT::U8x2, U8x2, PT::U16x2, U16x2),
(PT::U8x3, U8x3, PT::U16x3, U16x3),
(PT::U8x4, U8x4, PT::U16x4, U16x4),
)
}
/// Mapping in the forward direction of pixel's components of source image
/// into corresponding components of destination image.
pub fn forward_map(
&self,
src_image: &DynamicImageView,
dst_image: &mut DynamicImageViewMut,
src_image: &impl IntoImageView,
dst_image: &mut impl IntoImageViewMut,
) -> Result<(), MappingError> {
Self::map(&self.forward_mapping_tables, src_image, dst_image)
}
pub fn forward_map_inplace(&self, image: &mut DynamicImageViewMut) -> Result<(), MappingError> {
pub fn forward_map_inplace(
&self,
image: &mut impl IntoImageViewMut,
) -> Result<(), MappingError> {
Self::map_inplace(&self.forward_mapping_tables, image)
}
@@ -287,15 +346,15 @@ impl PixelComponentMapper {
/// into corresponding components of destination image.
pub fn backward_map(
&self,
src_image: &DynamicImageView,
dst_image: &mut DynamicImageViewMut,
src_image: &impl IntoImageView,
dst_image: &mut impl IntoImageViewMut,
) -> Result<(), MappingError> {
Self::map(&self.backward_mapping_tables, src_image, dst_image)
}
pub fn backward_map_inplace(
&self,
image: &mut DynamicImageViewMut,
image: &mut impl IntoImageViewMut,
) -> Result<(), MappingError> {
Self::map_inplace(&self.backward_mapping_tables, image)
}
+7 -7
View File
@@ -1,5 +1,5 @@
use crate::cpu_extensions::CpuExtensions;
use crate::pixels::F32;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use super::{Coefficients, Convolution};
@@ -8,22 +8,22 @@ mod native;
impl Convolution for F32 {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
_cpu_extensions: CpuExtensions,
) {
native::horiz_convolution(src_image, dst_image, offset, coeffs);
native::horiz_convolution(src_view, dst_view, offset, coeffs);
}
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
_cpu_extensions: CpuExtensions,
) {
native::vert_convolution(src_image, dst_image, offset, coeffs);
native::vert_convolution(src_view, dst_view, offset, coeffs);
}
}
+8 -8
View File
@@ -3,14 +3,14 @@ use crate::pixels::F32;
use crate::{ImageView, ImageViewMut};
pub(crate) fn horiz_convolution(
src_image: &ImageView<F32>,
dst_image: &mut ImageViewMut<F32>,
src_view: &impl ImageView<Pixel = F32>,
dst_view: &mut impl ImageViewMut<Pixel = F32>,
offset: u32,
coeffs: Coefficients,
) {
let coefficients_chunks = coeffs.get_chunks();
let src_rows = src_image.iter_rows(offset);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(offset);
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, src_row) in dst_rows.zip(src_rows) {
for (dst_pixel, coeffs_chunk) in dst_row.iter_mut().zip(&coefficients_chunks) {
let first_x_src = coeffs_chunk.start as usize;
@@ -25,20 +25,20 @@ pub(crate) fn horiz_convolution(
}
pub(crate) fn vert_convolution(
src_image: &ImageView<F32>,
dst_image: &mut ImageViewMut<F32>,
src_view: &impl ImageView<Pixel = F32>,
dst_view: &mut impl ImageViewMut<Pixel = F32>,
offset: u32,
coeffs: Coefficients,
) {
let coefficients_chunks = coeffs.get_chunks();
let dst_rows = dst_image.iter_rows_mut();
let dst_rows = dst_view.iter_rows_mut(0);
let start_src_x = offset as usize;
for (&coeffs_chunk, dst_row) in coefficients_chunks.iter().zip(dst_rows) {
let first_y_src = coeffs_chunk.start;
let mut src_x = start_src_x;
for dst_pixel in dst_row.iter_mut() {
let mut ss = 0.;
let src_rows = src_image.iter_rows(first_y_src);
let src_rows = src_view.iter_rows(first_y_src);
for (src_row, &k) in src_rows.zip(coeffs_chunk.values) {
let src_pixel = unsafe { src_row.get_unchecked(src_x) };
ss += src_pixel.0 as f64 * k;
+7 -8
View File
@@ -1,6 +1,5 @@
use crate::pixels::I32;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageView, ImageViewMut};
use super::{Coefficients, Convolution};
@@ -8,22 +7,22 @@ mod native;
impl Convolution for I32 {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
_cpu_extensions: CpuExtensions,
) {
native::horiz_convolution(src_image, dst_image, offset, coeffs);
native::horiz_convolution(src_view, dst_view, offset, coeffs);
}
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
_cpu_extensions: CpuExtensions,
) {
native::vert_convolution(src_image, dst_image, offset, coeffs);
native::vert_convolution(src_view, dst_view, offset, coeffs);
}
}
+8 -8
View File
@@ -3,14 +3,14 @@ use crate::pixels::I32;
use crate::{ImageView, ImageViewMut};
pub(crate) fn horiz_convolution(
src_image: &ImageView<I32>,
dst_image: &mut ImageViewMut<I32>,
src_view: &impl ImageView<Pixel = I32>,
dst_view: &mut impl ImageViewMut<Pixel = I32>,
offset: u32,
coeffs: Coefficients,
) {
let coefficients_chunks = coeffs.get_chunks();
let src_rows = src_image.iter_rows(offset);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(offset);
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, src_row) in dst_rows.zip(src_rows) {
for (dst_pixel, coeffs_chunk) in dst_row.iter_mut().zip(&coefficients_chunks) {
let first_x_src = coeffs_chunk.start as usize;
@@ -25,20 +25,20 @@ pub(crate) fn horiz_convolution(
}
pub(crate) fn vert_convolution(
src_image: &ImageView<I32>,
dst_image: &mut ImageViewMut<I32>,
src_view: &impl ImageView<Pixel = I32>,
dst_view: &mut impl ImageViewMut<Pixel = I32>,
offset: u32,
coeffs: Coefficients,
) {
let coefficients_chunks = coeffs.get_chunks();
let dst_rows = dst_image.iter_rows_mut();
let dst_rows = dst_view.iter_rows_mut(0);
let start_src_x = offset as usize;
for (&coeffs_chunk, dst_row) in coefficients_chunks.iter().zip(dst_rows) {
let first_y_src = coeffs_chunk.start;
let mut src_x = start_src_x;
for dst_pixel in dst_row.iter_mut() {
let mut ss = 0.;
let src_rows = src_image.iter_rows(first_y_src);
let src_rows = src_view.iter_rows(first_y_src);
for (src_row, &k) in src_rows.zip(coeffs_chunk.values) {
let src_pixel = unsafe { src_row.get_unchecked(src_x) };
ss += src_pixel.0 as f64 * k;
+16 -19
View File
@@ -1,10 +1,7 @@
use std::num::NonZeroU32;
pub use filters::*;
use crate::pixels::PixelExt;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::pixels::InnerPixel;
use crate::{CpuExtensions, ImageView, ImageViewMut};
#[macro_use]
mod macros;
@@ -28,21 +25,18 @@ cfg_if::cfg_if! {
}
}
pub(crate) trait Convolution
where
Self: PixelExt,
{
pub(crate) trait Convolution: InnerPixel {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
);
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
@@ -55,7 +49,7 @@ pub(crate) struct Bound {
pub size: u32,
}
#[derive(Debug, Clone)]
#[derive(Debug, Clone, Default)]
pub(crate) struct Coefficients {
pub values: Vec<f64>,
pub window_size: usize,
@@ -86,17 +80,20 @@ impl Coefficients {
}
pub(crate) fn precompute_coefficients(
in_size: NonZeroU32,
in_size: u32,
in0: f64, // Left border for cropping
in1: f64, // Right border for cropping
out_size: NonZeroU32,
out_size: u32,
filter: fn(f64) -> f64,
filter_support: f64,
) -> Coefficients {
let in_size = in_size.get();
let out_size = out_size.get();
if in_size == 0 || out_size == 0 {
return Coefficients::default();
}
let scale = (in1 - in0) / out_size as f64;
if scale <= 0. {
return Coefficients::default();
}
let filter_scale = scale.max(1.0);
// Determine filter radius size (length of resampling filter)
+12 -17
View File
@@ -2,39 +2,34 @@ use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U16;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16>,
dst_image: &mut ImageViewMut<U16>,
src_view: &impl ImageView<Pixel = U16>,
dst_view: &mut impl ImageViewMut<Pixel = U16>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -46,7 +41,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "avx2")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U16]; 4],
dst_rows: [&mut &mut [U16]; 4],
dst_rows: [&mut [U16]; 4],
coefficients_chunks: &[optimisations::CoefficientsI32Chunk],
normalizer: &optimisations::Normalizer32,
) {
+11 -14
View File
@@ -1,7 +1,6 @@
use crate::convolution::vertical_u16::vert_convolution_u16;
use crate::pixels::U16;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageView, ImageViewMut};
use super::{Coefficients, Convolution};
@@ -17,34 +16,32 @@ mod wasm32;
impl Convolution for U16 {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => avx2::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Avx2 => avx2::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => neon::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Neon => neon::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => {
wasm32::horiz_convolution(src_image, dst_image, offset, coeffs)
}
_ => native::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Simd128 => wasm32::horiz_convolution(src_view, dst_view, offset, coeffs),
_ => native::horiz_convolution(src_view, dst_view, offset, coeffs),
}
}
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
vert_convolution_u16(src_image, dst_image, offset, coeffs, cpu_extensions);
vert_convolution_u16(src_view, dst_view, offset, coeffs, cpu_extensions);
}
}
+4 -4
View File
@@ -4,8 +4,8 @@ use crate::{ImageView, ImageViewMut};
#[inline(always)]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16>,
dst_image: &mut ImageViewMut<U16>,
src_view: &impl ImageView<Pixel = U16>,
dst_view: &mut impl ImageViewMut<Pixel = U16>,
offset: u32,
coeffs: Coefficients,
) {
@@ -14,8 +14,8 @@ pub(crate) fn horiz_convolution(
let coefficients_chunks = normalizer.normalized_chunks();
let initial = 1i64 << (precision - 1);
let src_rows = src_image.iter_rows(offset);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(offset);
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, src_row) in dst_rows.zip(src_rows) {
for (&coeffs_chunk, dst_pixel) in coefficients_chunks.iter().zip(dst_row.iter_mut()) {
let first_x_src = coeffs_chunk.start as usize;
+12 -17
View File
@@ -2,39 +2,34 @@ use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U16;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16>,
dst_image: &mut ImageViewMut<U16>,
src_view: &impl ImageView<Pixel = U16>,
dst_view: &mut impl ImageViewMut<Pixel = U16>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -46,7 +41,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "sse4.1")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U16]; 4],
dst_rows: [&mut &mut [U16]; 4],
dst_rows: [&mut [U16]; 4],
coefficients_chunks: &[optimisations::CoefficientsI32Chunk],
normalizer: &optimisations::Normalizer32,
) {
+12 -17
View File
@@ -2,39 +2,34 @@ use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U16x2;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16x2>,
dst_image: &mut ImageViewMut<U16x2>,
src_view: &impl ImageView<Pixel = U16x2>,
dst_view: &mut impl ImageViewMut<Pixel = U16x2>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -46,7 +41,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "avx2")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U16x2]; 4],
dst_rows: [&mut &mut [U16x2]; 4],
dst_rows: [&mut [U16x2]; 4],
coefficients_chunks: &[optimisations::CoefficientsI32Chunk],
normalizer: &optimisations::Normalizer32,
) {
+11 -14
View File
@@ -1,7 +1,6 @@
use crate::convolution::vertical_u16::vert_convolution_u16;
use crate::pixels::U16x2;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageView, ImageViewMut};
use super::{Coefficients, Convolution};
@@ -17,34 +16,32 @@ mod wasm32;
impl Convolution for U16x2 {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => avx2::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Avx2 => avx2::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => neon::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Neon => neon::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => {
wasm32::horiz_convolution(src_image, dst_image, offset, coeffs)
}
_ => native::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Simd128 => wasm32::horiz_convolution(src_view, dst_view, offset, coeffs),
_ => native::horiz_convolution(src_view, dst_view, offset, coeffs),
}
}
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
vert_convolution_u16(src_image, dst_image, offset, coeffs, cpu_extensions);
vert_convolution_u16(src_view, dst_view, offset, coeffs, cpu_extensions);
}
}
+4 -4
View File
@@ -4,8 +4,8 @@ use crate::{ImageView, ImageViewMut};
#[inline(always)]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16x2>,
dst_image: &mut ImageViewMut<U16x2>,
src_view: &impl ImageView<Pixel = U16x2>,
dst_view: &mut impl ImageViewMut<Pixel = U16x2>,
offset: u32,
coeffs: Coefficients,
) {
@@ -14,8 +14,8 @@ pub(crate) fn horiz_convolution(
let coefficients_chunks = normalizer.normalized_chunks();
let initial: i64 = 1 << (precision - 1);
let src_rows = src_image.iter_rows(offset);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(offset);
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, src_row) in dst_rows.zip(src_rows) {
for (&coeffs_chunk, dst_pixel) in coefficients_chunks.iter().zip(dst_row.iter_mut()) {
let first_x_src = coeffs_chunk.start as usize;
+12 -17
View File
@@ -2,39 +2,34 @@ use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U16x2;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16x2>,
dst_image: &mut ImageViewMut<U16x2>,
src_view: &impl ImageView<Pixel = U16x2>,
dst_view: &mut impl ImageViewMut<Pixel = U16x2>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -47,7 +42,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "sse4.1")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U16x2]; 4],
dst_rows: [&mut &mut [U16x2]; 4],
dst_rows: [&mut [U16x2]; 4],
coefficients_chunks: &[optimisations::CoefficientsI32Chunk],
normalizer: &optimisations::Normalizer32,
) {
+12 -17
View File
@@ -1,40 +1,35 @@
use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::image_view::{ImageView, ImageViewMut};
use crate::pixels::U16x3;
use crate::simd_utils;
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16x3>,
dst_image: &mut ImageViewMut<U16x3>,
src_view: &impl ImageView<Pixel = U16x3>,
dst_view: &mut impl ImageViewMut<Pixel = U16x3>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -46,7 +41,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "avx2")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U16x3]; 4],
dst_rows: [&mut &mut [U16x3]; 4],
dst_rows: [&mut [U16x3]; 4],
coefficients_chunks: &[optimisations::CoefficientsI32Chunk],
normalizer: &optimisations::Normalizer32,
) {
+11 -14
View File
@@ -1,7 +1,6 @@
use crate::convolution::vertical_u16::vert_convolution_u16;
use crate::pixels::U16x3;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageView, ImageViewMut};
use super::{Coefficients, Convolution};
@@ -17,34 +16,32 @@ mod wasm32;
impl Convolution for U16x3 {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => avx2::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Avx2 => avx2::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => neon::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Neon => neon::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => {
wasm32::horiz_convolution(src_image, dst_image, offset, coeffs)
}
_ => native::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Simd128 => wasm32::horiz_convolution(src_view, dst_view, offset, coeffs),
_ => native::horiz_convolution(src_view, dst_view, offset, coeffs),
}
}
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
vert_convolution_u16(src_image, dst_image, offset, coeffs, cpu_extensions);
vert_convolution_u16(src_view, dst_view, offset, coeffs, cpu_extensions);
}
}
+6 -6
View File
@@ -4,8 +4,8 @@ use crate::{ImageView, ImageViewMut};
#[inline(always)]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16x3>,
dst_image: &mut ImageViewMut<U16x3>,
src_view: &impl ImageView<Pixel = U16x3>,
dst_view: &mut impl ImageViewMut<Pixel = U16x3>,
offset: u32,
coeffs: Coefficients,
) {
@@ -14,16 +14,16 @@ pub(crate) fn horiz_convolution(
let coefficients_chunks = normalizer.normalized_chunks();
let initial = 1i64 << (precision - 1);
let src_rows = src_image.iter_rows(offset);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(offset);
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, src_row) in dst_rows.zip(src_rows) {
for (&coeffs_chunk, dst_pixel) in coefficients_chunks.iter().zip(dst_row.iter_mut()) {
let first_x_src = coeffs_chunk.start as usize;
let mut ss = [initial; 3];
let src_pixels = unsafe { src_row.get_unchecked(first_x_src..) };
for (&k, src_pixel) in coeffs_chunk.values.iter().zip(src_pixels) {
for (i, s) in ss.iter_mut().enumerate() {
*s += src_pixel.0[i] as i64 * (k as i64);
for (s, c) in ss.iter_mut().zip(src_pixel.0) {
*s += c as i64 * (k as i64);
}
}
for (i, s) in ss.iter().copied().enumerate() {
+17 -23
View File
@@ -1,41 +1,35 @@
use std::arch::x86_64::*;
use crate::convolution::optimisations::CoefficientsI32Chunk;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U16x3;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16x3>,
dst_image: &mut ImageViewMut<U16x3>,
src_view: &impl ImageView<Pixel = U16x3>,
dst_view: &mut impl ImageViewMut<Pixel = U16x3>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_8u4x(src_rows, dst_rows, &coefficients_chunks, &normalizer);
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_8u(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -46,10 +40,10 @@ pub(crate) fn horiz_convolution(
/// - max(chunk.start + chunk.values.len() for chunk in coefficients_chunks) <= src_row.0.len()
/// - precision <= MAX_COEFS_PRECISION
#[target_feature(enable = "sse4.1")]
unsafe fn horiz_convolution_8u4x(
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U16x3]; 4],
dst_rows: [&mut &mut [U16x3]; 4],
coefficients_chunks: &[CoefficientsI32Chunk],
dst_rows: [&mut [U16x3]; 4],
coefficients_chunks: &[optimisations::CoefficientsI32Chunk],
normalizer: &optimisations::Normalizer32,
) {
let precision = normalizer.precision();
@@ -141,10 +135,10 @@ unsafe fn horiz_convolution_8u4x(
/// - max(chunk.start + chunk.values.len() for chunk in coefficients_chunks) <= src_row.len()
/// - precision <= MAX_COEFS_PRECISION
#[target_feature(enable = "sse4.1")]
unsafe fn horiz_convolution_8u(
unsafe fn horiz_convolution_one_row(
src_row: &[U16x3],
dst_row: &mut [U16x3],
coefficients_chunks: &[CoefficientsI32Chunk],
coefficients_chunks: &[optimisations::CoefficientsI32Chunk],
normalizer: &optimisations::Normalizer32,
) {
let precision = normalizer.precision();
+12 -17
View File
@@ -2,39 +2,34 @@ use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U16x4;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16x4>,
dst_image: &mut ImageViewMut<U16x4>,
src_view: &impl ImageView<Pixel = U16x4>,
dst_view: &mut impl ImageViewMut<Pixel = U16x4>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -47,7 +42,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "avx2")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U16x4]; 4],
dst_rows: [&mut &mut [U16x4]; 4],
dst_rows: [&mut [U16x4]; 4],
coefficients_chunks: &[optimisations::CoefficientsI32Chunk],
normalizer: &optimisations::Normalizer32,
) {
+11 -14
View File
@@ -1,7 +1,6 @@
use crate::convolution::vertical_u16::vert_convolution_u16;
use crate::pixels::U16x4;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageView, ImageViewMut};
use super::{Coefficients, Convolution};
@@ -17,34 +16,32 @@ mod wasm32;
impl Convolution for U16x4 {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => avx2::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Avx2 => avx2::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => neon::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Neon => neon::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => {
wasm32::horiz_convolution(src_image, dst_image, offset, coeffs)
}
_ => native::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Simd128 => wasm32::horiz_convolution(src_view, dst_view, offset, coeffs),
_ => native::horiz_convolution(src_view, dst_view, offset, coeffs),
}
}
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
vert_convolution_u16(src_image, dst_image, offset, coeffs, cpu_extensions);
vert_convolution_u16(src_view, dst_view, offset, coeffs, cpu_extensions);
}
}
+4 -4
View File
@@ -4,8 +4,8 @@ use crate::{ImageView, ImageViewMut};
#[inline(always)]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16x4>,
dst_image: &mut ImageViewMut<U16x4>,
src_view: &impl ImageView<Pixel = U16x4>,
dst_view: &mut impl ImageViewMut<Pixel = U16x4>,
offset: u32,
coeffs: Coefficients,
) {
@@ -14,8 +14,8 @@ pub(crate) fn horiz_convolution(
let coefficients_chunks = normalizer.normalized_chunks();
let initial: i64 = 1 << (precision - 1);
let src_rows = src_image.iter_rows(offset);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(offset);
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, src_row) in dst_rows.zip(src_rows) {
for (&coeffs_chunk, dst_pixel) in coefficients_chunks.iter().zip(dst_row.iter_mut()) {
let first_x_src = coeffs_chunk.start as usize;
+12 -17
View File
@@ -2,39 +2,34 @@ use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U16x4;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U16x4>,
dst_image: &mut ImageViewMut<U16x4>,
src_view: &impl ImageView<Pixel = U16x4>,
dst_view: &mut impl ImageViewMut<Pixel = U16x4>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -47,7 +42,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "sse4.1")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U16x4]; 4],
dst_rows: [&mut &mut [U16x4]; 4],
dst_rows: [&mut [U16x4]; 4],
coefficients_chunks: &[optimisations::CoefficientsI32Chunk],
normalizer: &optimisations::Normalizer32,
) {
+12 -17
View File
@@ -2,39 +2,34 @@ use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U8;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8>,
dst_image: &mut ImageViewMut<U8>,
src_view: &impl ImageView<Pixel = U8>,
dst_view: &mut impl ImageViewMut<Pixel = U8>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer16::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -48,7 +43,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "avx2")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U8]; 4],
dst_rows: [&mut &mut [U8]; 4],
dst_rows: [&mut [U8]; 4],
coefficients_chunks: &[optimisations::CoefficientsI16Chunk],
normalizer: &optimisations::Normalizer16,
) {
+11 -14
View File
@@ -1,7 +1,6 @@
use crate::convolution::vertical_u8::vert_convolution_u8;
use crate::pixels::U8;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageView, ImageViewMut};
use super::{Coefficients, Convolution};
@@ -17,34 +16,32 @@ mod wasm32;
impl Convolution for U8 {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => avx2::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Avx2 => avx2::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => neon::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Neon => neon::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => {
wasm32::horiz_convolution(src_image, dst_image, offset, coeffs)
}
_ => native::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Simd128 => wasm32::horiz_convolution(src_view, dst_view, offset, coeffs),
_ => native::horiz_convolution(src_view, dst_view, offset, coeffs),
}
}
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
vert_convolution_u8(src_image, dst_image, offset, coeffs, cpu_extensions);
vert_convolution_u8(src_view, dst_view, offset, coeffs, cpu_extensions);
}
}
+4 -4
View File
@@ -4,8 +4,8 @@ use crate::{ImageView, ImageViewMut};
#[inline(always)]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8>,
dst_image: &mut ImageViewMut<U8>,
src_view: &impl ImageView<Pixel = U8>,
dst_view: &mut impl ImageViewMut<Pixel = U8>,
offset: u32,
coeffs: Coefficients,
) {
@@ -14,8 +14,8 @@ pub(crate) fn horiz_convolution(
let coefficients_chunks = normalizer.normalized_chunks();
let initial = 1i32 << (precision - 1);
let src_rows = src_image.iter_rows(offset);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(offset);
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, src_row) in dst_rows.zip(src_rows) {
for (&coeffs_chunk, dst_pixel) in coefficients_chunks.iter().zip(dst_row.iter_mut()) {
let first_x_src = coeffs_chunk.start as usize;
+12 -17
View File
@@ -2,39 +2,34 @@ use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U8;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8>,
dst_image: &mut ImageViewMut<U8>,
src_view: &impl ImageView<Pixel = U8>,
dst_view: &mut impl ImageViewMut<Pixel = U8>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer16::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -48,7 +43,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "sse4.1")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U8]; 4],
dst_rows: [&mut &mut [U8]; 4],
dst_rows: [&mut [U8]; 4],
coefficients_chunks: &[optimisations::CoefficientsI16Chunk],
normalizer: &optimisations::Normalizer16,
) {
+22 -32
View File
@@ -1,45 +1,35 @@
use std::arch::x86_64::*;
use crate::convolution::{Coefficients, optimisations};
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U8x2;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8x2>,
dst_image: &mut ImageViewMut<U8x2>,
src_view: &impl ImageView<Pixel = U8x2>,
dst_view: &mut impl ImageViewMut<Pixel = U8x2>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer16::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(
src_rows,
dst_rows,
&coefficients_chunks,
&normalizer,
);
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -53,7 +43,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "avx2")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U8x2]; 4],
dst_rows: [&mut &mut [U8x2]; 4],
dst_rows: [&mut [U8x2]; 4],
coefficients_chunks: &[optimisations::CoefficientsI16Chunk],
normalizer: &optimisations::Normalizer16,
) {
@@ -201,7 +191,7 @@ unsafe fn horiz_convolution_four_rows(
#[target_feature(enable = "avx2")]
unsafe fn set_dst_pixel(
raw: __m128i,
d_row: &mut &mut [U8x2],
d_row: &mut [U8x2],
dst_x: usize,
normalizer: &optimisations::Normalizer16,
) {
@@ -211,7 +201,7 @@ unsafe fn set_dst_pixel(
let a32 = ((a32x2 >> 32) as i32).saturating_add((a32x2 & 0xffffffff) as i32);
let l8 = normalizer.clip(l32);
let a8 = normalizer.clip(a32);
d_row.get_unchecked_mut(dst_x).0 = u16::from_le_bytes([l8, a8]);
d_row.get_unchecked_mut(dst_x).0 = [l8, a8];
}
/// For safety, it is necessary to ensure the following conditions:
@@ -257,8 +247,8 @@ unsafe fn horiz_convolution_one_row(
*/
#[rustfmt::skip]
let coeff_sh1 = _mm256_set_epi8(
7, 6, 5, 4, 7, 6, 5, 4, 3, 2, 1, 0, 3,2, 1, 0,
7, 6, 5, 4, 7, 6, 5, 4, 3, 2, 1, 0, 3,2, 1, 0,
7, 6, 5, 4, 7, 6, 5, 4, 3, 2, 1, 0, 3, 2, 1, 0,
7, 6, 5, 4, 7, 6, 5, 4, 3, 2, 1, 0, 3, 2, 1, 0,
);
/*
@@ -275,7 +265,7 @@ unsafe fn horiz_convolution_one_row(
#[rustfmt::skip]
let pix_sh2 = _mm256_set_epi8(
-1, 15, -1, 13, -1, 14, -1, 12, -1, 11, -1, 9, -1, 10, -1, 8,
-1, 15, -1, 13, -1, 14, -1, 12, -1, 11, -1, 9, -1, 10, -1, 8,
-1, 15, -1, 13, -1, 14, -1, 12, -1, 11, -1, 9, -1, 10, -1, 8,
);
/*
|C0 | |C1 | |C2 | |C3 | |C4 | |C5 | |C6 | |C7 |
@@ -310,7 +300,7 @@ unsafe fn horiz_convolution_one_row(
#[rustfmt::skip]
let coeff_sh3 = _mm256_set_epi8(
15, 14, 13, 12, 15, 14, 13, 12, 11, 10, 9, 8, 11, 10, 9, 8,
7, 6, 5, 4, 7, 6, 5, 4, 3, 2, 1, 0, 3,2, 1, 0,
7, 6, 5, 4, 7, 6, 5, 4, 3, 2, 1, 0, 3, 2, 1, 0,
);
/*
@@ -395,7 +385,7 @@ unsafe fn horiz_convolution_one_row(
let mut coeffs: [i16; 3] = [0; 3];
for (i, &coeff) in reminder1.iter().enumerate() {
coeffs[i] = coeff;
let pixel: [u8; 2] = src_row.get_unchecked(x).0.to_le_bytes();
let pixel: [u8; 2] = src_row.get_unchecked(x).0;
pixels[i * 2] = pixel[0] as i16;
pixels[i * 2 + 1] = pixel[1] as i16;
x += 1;
@@ -418,6 +408,6 @@ unsafe fn horiz_convolution_one_row(
let l32 = ((lo & 0xffffffff) as i32).saturating_add((hi & 0xffffffff) as i32);
let a8 = normalizer.clip(a32);
let l8 = normalizer.clip(l32);
dst_row.get_unchecked_mut(dst_x).0 = u16::from_le_bytes([l8, a8]);
dst_row.get_unchecked_mut(dst_x).0 = [l8, a8];
}
}
+11 -14
View File
@@ -1,7 +1,6 @@
use crate::convolution::vertical_u8::vert_convolution_u8;
use crate::pixels::U8x2;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageView, ImageViewMut};
use super::{Coefficients, Convolution};
@@ -17,34 +16,32 @@ mod wasm32;
impl Convolution for U8x2 {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => avx2::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Avx2 => avx2::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => neon::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Neon => neon::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => {
wasm32::horiz_convolution(src_image, dst_image, offset, coeffs)
}
_ => native::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Simd128 => wasm32::horiz_convolution(src_view, dst_view, offset, coeffs),
_ => native::horiz_convolution(src_view, dst_view, offset, coeffs),
}
}
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
vert_convolution_u8(src_image, dst_image, offset, coeffs, cpu_extensions);
vert_convolution_u8(src_view, dst_view, offset, coeffs, cpu_extensions);
}
}
+7 -8
View File
@@ -3,8 +3,8 @@ use crate::pixels::U8x2;
use crate::{ImageView, ImageViewMut};
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8x2>,
dst_image: &mut ImageViewMut<U8x2>,
src_view: &impl ImageView<Pixel = U8x2>,
dst_view: &mut impl ImageViewMut<Pixel = U8x2>,
offset: u32,
coeffs: Coefficients,
) {
@@ -13,8 +13,8 @@ pub(crate) fn horiz_convolution(
let coefficients_chunks = normalizer.normalized_chunks();
let initial = 1 << (precision - 1);
let src_rows = src_image.iter_rows(offset);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(offset);
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, src_row) in dst_rows.zip(src_rows) {
for (&coeffs_chunk, dst_pixel) in coefficients_chunks.iter().zip(dst_row.iter_mut()) {
let first_x_src = coeffs_chunk.start as usize;
@@ -22,12 +22,11 @@ pub(crate) fn horiz_convolution(
let mut ss = [initial; 2];
let src_pixels = unsafe { src_row.get_unchecked(first_x_src..) };
for (&k, &src_pixel) in ks.iter().zip(src_pixels) {
let components: [u8; 2] = src_pixel.0.to_le_bytes();
for (i, s) in ss.iter_mut().enumerate() {
*s += components[i] as i32 * (k as i32);
for (s, c) in ss.iter_mut().zip(src_pixel.0) {
*s += c as i32 * (k as i32);
}
}
dst_pixel.0 = u16::from_le_bytes(ss.map(|v| unsafe { normalizer.clip(v) }));
dst_pixel.0 = ss.map(|v| unsafe { normalizer.clip(v) });
}
}
}
+10 -15
View File
@@ -1,14 +1,13 @@
use std::arch::aarch64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::neon_utils;
use crate::pixels::U8x2;
use crate::{ImageView, ImageViewMut};
use crate::{neon_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8x2>,
dst_image: &mut ImageViewMut<U8x2>,
src_view: &impl ImageView<Pixel = U8x2>,
dst_view: &mut impl ImageViewMut<Pixel = U8x2>,
offset: u32,
coeffs: Coefficients,
) {
@@ -17,8 +16,8 @@ pub(crate) fn horiz_convolution(
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, precision);
@@ -26,16 +25,12 @@ pub(crate) fn horiz_convolution(
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
precision,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -48,7 +43,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "neon")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U8x2]; 4],
dst_rows: [&mut &mut [U8x2]; 4],
dst_rows: [&mut [U8x2]; 4],
coefficients_chunks: &[optimisations::CoefficientsI16Chunk],
precision: u8,
) {
+17 -22
View File
@@ -2,39 +2,34 @@ use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U8x2;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8x2>,
dst_image: &mut ImageViewMut<U8x2>,
src_view: &impl ImageView<Pixel = U8x2>,
dst_view: &mut impl ImageViewMut<Pixel = U8x2>,
offset: u32,
coeffs: Coefficients,
) {
let normalizer = optimisations::Normalizer16::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows(src_rows, dst_rows, &coefficients_chunks, &normalizer);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
&normalizer,
);
horiz_convolution_one_row(src_row, dst_row, &coefficients_chunks, &normalizer);
}
yy += 1;
}
}
@@ -48,7 +43,7 @@ pub(crate) fn horiz_convolution(
#[target_feature(enable = "sse4.1")]
unsafe fn horiz_convolution_four_rows(
src_rows: [&[U8x2]; 4],
dst_rows: [&mut &mut [U8x2]; 4],
dst_rows: [&mut [U8x2]; 4],
coefficients_chunks: &[optimisations::CoefficientsI16Chunk],
normalizer: &optimisations::Normalizer16,
) {
@@ -148,7 +143,7 @@ unsafe fn horiz_convolution_four_rows(
#[target_feature(enable = "sse4.1")]
unsafe fn set_dst_pixel(
raw: __m128i,
d_row: &mut &mut [U8x2],
d_row: &mut [U8x2],
dst_x: usize,
normalizer: &optimisations::Normalizer16,
) {
@@ -158,7 +153,7 @@ unsafe fn set_dst_pixel(
let a32 = ((a32x2 >> 32) as i32).saturating_add((a32x2 & 0xffffffff) as i32);
let l8 = normalizer.clip(l32);
let a8 = normalizer.clip(a32);
d_row.get_unchecked_mut(dst_x).0 = u16::from_le_bytes([l8, a8]);
d_row.get_unchecked_mut(dst_x).0 = [l8, a8];
}
/// For safety, it is necessary to ensure the following conditions:
@@ -203,7 +198,7 @@ unsafe fn horiz_convolution_one_row(
*/
#[rustfmt::skip]
let coeff_sh1 = _mm_set_epi8(
7, 6, 5, 4, 7, 6, 5, 4, 3, 2, 1, 0, 3,2, 1, 0,
7, 6, 5, 4, 7, 6, 5, 4, 3, 2, 1, 0, 3, 2, 1, 0,
);
/*
@@ -292,7 +287,7 @@ unsafe fn horiz_convolution_one_row(
let mut coeffs: [i16; 3] = [0; 3];
for (i, &coeff) in reminder1.iter().enumerate() {
coeffs[i] = coeff;
let pixel: [u8; 2] = src_row.get_unchecked(x).0.to_le_bytes();
let pixel: [u8; 2] = src_row.get_unchecked(x).0;
pixels[i * 2] = pixel[0] as i16;
pixels[i * 2 + 1] = pixel[1] as i16;
x += 1;
@@ -315,6 +310,6 @@ unsafe fn horiz_convolution_one_row(
let l32 = ((lo & 0xffffffff) as i32).saturating_add((hi & 0xffffffff) as i32);
let a8 = normalizer.clip(a32);
let l8 = normalizer.clip(l32);
dst_row.get_unchecked_mut(dst_x).0 = u16::from_le_bytes([l8, a8]);
dst_row.get_unchecked_mut(dst_x).0 = [l8, a8];
}
}
+15 -19
View File
@@ -3,13 +3,12 @@ use std::intrinsics::transmute;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U8x3;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8x3>,
dst_image: &mut ImageViewMut<U8x3>,
src_view: &impl ImageView<Pixel = U8x3>,
dst_view: &mut impl ImageViewMut<Pixel = U8x3>,
offset: u32,
coeffs: Coefficients,
) {
@@ -18,39 +17,36 @@ pub(crate) fn horiz_convolution(
macro_rules! call {
($imm8:expr) => {{
horiz_convolution_p::<$imm8>(src_image, dst_image, offset, normalizer);
horiz_convolution_p::<$imm8>(src_view, dst_view, offset, normalizer);
}};
}
constify_imm8!(precision, call);
}
fn horiz_convolution_p<const PRECISION: i32>(
src_image: &ImageView<U8x3>,
dst_image: &mut ImageViewMut<U8x3>,
src_view: &impl ImageView<Pixel = U8x3>,
dst_view: &mut impl ImageViewMut<Pixel = U8x3>,
offset: u32,
normalizer: optimisations::Normalizer16,
) {
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows::<PRECISION>(src_rows, dst_rows, &coefficients_chunks);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row::<PRECISION>(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
);
horiz_convolution_one_row::<PRECISION>(src_row, dst_row, &coefficients_chunks);
}
yy += 1;
}
}
@@ -64,7 +60,7 @@ fn horiz_convolution_p<const PRECISION: i32>(
#[target_feature(enable = "avx2")]
unsafe fn horiz_convolution_four_rows<const PRECISION: i32>(
src_rows: [&[U8x3]; 4],
dst_rows: [&mut &mut [U8x3]; 4],
dst_rows: [&mut [U8x3]; 4],
coefficients_chunks: &[optimisations::CoefficientsI16Chunk],
) {
let zero = _mm256_setzero_si256();
+11 -14
View File
@@ -1,7 +1,6 @@
use crate::convolution::vertical_u8::vert_convolution_u8;
use crate::pixels::U8x3;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageView, ImageViewMut};
use super::{Coefficients, Convolution};
@@ -17,34 +16,32 @@ mod wasm32;
impl Convolution for U8x3 {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => avx2::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Avx2 => avx2::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => neon::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Neon => neon::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => {
wasm32::horiz_convolution(src_image, dst_image, offset, coeffs)
}
_ => native::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Simd128 => wasm32::horiz_convolution(src_view, dst_view, offset, coeffs),
_ => native::horiz_convolution(src_view, dst_view, offset, coeffs),
}
}
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
vert_convolution_u8(src_image, dst_image, offset, coeffs, cpu_extensions);
vert_convolution_u8(src_view, dst_view, offset, coeffs, cpu_extensions);
}
}
+7 -9
View File
@@ -4,8 +4,8 @@ use crate::{ImageView, ImageViewMut};
#[inline(always)]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8x3>,
dst_image: &mut ImageViewMut<U8x3>,
src_view: &impl ImageView<Pixel = U8x3>,
dst_view: &mut impl ImageViewMut<Pixel = U8x3>,
offset: u32,
coeffs: Coefficients,
) {
@@ -14,21 +14,19 @@ pub(crate) fn horiz_convolution(
let coefficients_chunks = normalizer.normalized_chunks();
let initial = 1i32 << (precision - 1);
let src_rows = src_image.iter_rows(offset);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(offset);
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, src_row) in dst_rows.zip(src_rows) {
for (&coeffs_chunk, dst_pixel) in coefficients_chunks.iter().zip(dst_row.iter_mut()) {
let first_x_src = coeffs_chunk.start as usize;
let mut ss = [initial; 3];
let src_pixels = unsafe { src_row.get_unchecked(first_x_src..) };
for (&k, src_pixel) in coeffs_chunk.values.iter().zip(src_pixels) {
for (i, s) in ss.iter_mut().enumerate() {
*s += src_pixel.0[i] as i32 * (k as i32);
for (s, c) in ss.iter_mut().zip(src_pixel.0) {
*s += c as i32 * (k as i32);
}
}
for (i, s) in ss.iter().copied().enumerate() {
dst_pixel.0[i] = unsafe { normalizer.clip(s) };
}
dst_pixel.0 = ss.map(|v| unsafe { normalizer.clip(v) });
}
}
}
+15 -19
View File
@@ -3,13 +3,12 @@ use std::intrinsics::transmute;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U8x3;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8x3>,
dst_image: &mut ImageViewMut<U8x3>,
src_view: &impl ImageView<Pixel = U8x3>,
dst_view: &mut impl ImageViewMut<Pixel = U8x3>,
offset: u32,
coeffs: Coefficients,
) {
@@ -18,39 +17,36 @@ pub(crate) fn horiz_convolution(
macro_rules! call {
($imm8:expr) => {{
horiz_convolution_p::<$imm8>(src_image, dst_image, offset, normalizer);
horiz_convolution_p::<$imm8>(src_view, dst_view, offset, normalizer);
}};
}
constify_imm8!(precision, call);
}
fn horiz_convolution_p<const PRECISION: i32>(
src_image: &ImageView<U8x3>,
dst_image: &mut ImageViewMut<U8x3>,
src_view: &impl ImageView<Pixel = U8x3>,
dst_view: &mut impl ImageViewMut<Pixel = U8x3>,
offset: u32,
normalizer: optimisations::Normalizer16,
) {
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows::<PRECISION>(src_rows, dst_rows, &coefficients_chunks);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row::<PRECISION>(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
);
horiz_convolution_one_row::<PRECISION>(src_row, dst_row, &coefficients_chunks);
}
yy += 1;
}
}
@@ -64,7 +60,7 @@ fn horiz_convolution_p<const PRECISION: i32>(
#[target_feature(enable = "sse4.1")]
unsafe fn horiz_convolution_four_rows<const PRECISION: i32>(
src_rows: [&[U8x3]; 4],
dst_rows: [&mut &mut [U8x3]; 4],
dst_rows: [&mut [U8x3]; 4],
coefficients_chunks: &[optimisations::CoefficientsI16Chunk],
) {
let zero = _mm_setzero_si128();
+15 -19
View File
@@ -3,16 +3,15 @@ use std::intrinsics::transmute;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U8x4;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
// This code is based on C-implementation from Pillow-SIMD package for Python
// https://github.com/uploadcare/pillow-simd
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8x4>,
dst_image: &mut ImageViewMut<U8x4>,
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
offset: u32,
coeffs: Coefficients,
) {
@@ -21,39 +20,36 @@ pub(crate) fn horiz_convolution(
macro_rules! call {
($imm8:expr) => {{
horiz_convolution_p::<$imm8>(src_image, dst_image, offset, normalizer);
horiz_convolution_p::<$imm8>(src_view, dst_view, offset, normalizer);
}};
}
constify_imm8!(precision, call);
}
fn horiz_convolution_p<const PRECISION: i32>(
src_image: &ImageView<U8x4>,
dst_image: &mut ImageViewMut<U8x4>,
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
offset: u32,
normalizer: optimisations::Normalizer16,
) {
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows::<PRECISION>(src_rows, dst_rows, &coefficients_chunks);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row::<PRECISION>(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
);
horiz_convolution_one_row::<PRECISION>(src_row, dst_row, &coefficients_chunks);
}
yy += 1;
}
}
@@ -67,7 +63,7 @@ fn horiz_convolution_p<const PRECISION: i32>(
#[target_feature(enable = "avx2")]
unsafe fn horiz_convolution_four_rows<const PRECISION: i32>(
src_rows: [&[U8x4]; 4],
dst_rows: [&mut &mut [U8x4]; 4],
dst_rows: [&mut [U8x4]; 4],
coefficients_chunks: &[optimisations::CoefficientsI16Chunk],
) {
let zero = _mm256_setzero_si256();
+11 -14
View File
@@ -1,7 +1,6 @@
use crate::convolution::vertical_u8::vert_convolution_u8;
use crate::pixels::U8x4;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::{CpuExtensions, ImageView, ImageViewMut};
use super::{Coefficients, Convolution};
@@ -17,34 +16,32 @@ mod wasm32;
impl Convolution for U8x4 {
fn horiz_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => avx2::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Avx2 => avx2::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Sse4_1 => sse4::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => neon::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Neon => neon::horiz_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => {
wasm32::horiz_convolution(src_image, dst_image, offset, coeffs)
}
_ => native::horiz_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Simd128 => wasm32::horiz_convolution(src_view, dst_view, offset, coeffs),
_ => native::horiz_convolution(src_view, dst_view, offset, coeffs),
}
}
fn vert_convolution(
src_image: &ImageView<Self>,
dst_image: &mut ImageViewMut<Self>,
src_view: &impl ImageView<Pixel = Self>,
dst_view: &mut impl ImageViewMut<Pixel = Self>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
vert_convolution_u8(src_image, dst_image, offset, coeffs, cpu_extensions);
vert_convolution_u8(src_view, dst_view, offset, coeffs, cpu_extensions);
}
}
+6 -10
View File
@@ -1,11 +1,11 @@
use crate::convolution::{optimisations, Coefficients};
use crate::image_view::{ImageView, ImageViewMut};
use crate::pixels::U8x4;
use crate::{ImageView, ImageViewMut};
#[inline(always)]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8x4>,
dst_image: &mut ImageViewMut<U8x4>,
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
offset: u32,
coeffs: Coefficients,
) {
@@ -14,23 +14,19 @@ pub(crate) fn horiz_convolution(
let coefficients_chunks = normalizer.normalized_chunks();
let initial = 1 << (precision - 1);
let src_rows = src_image.iter_rows(offset);
let dst_rows = dst_image.iter_rows_mut();
let src_rows = src_view.iter_rows(offset);
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, src_row) in dst_rows.zip(src_rows) {
for (&coeffs_chunk, dst_pixel) in coefficients_chunks.iter().zip(dst_row.iter_mut()) {
let first_x_src = coeffs_chunk.start as usize;
let mut ss = [initial; 4];
let src_pixels = unsafe { src_row.get_unchecked(first_x_src..) };
for (&k, &src_pixel) in coeffs_chunk.values.iter().zip(src_pixels) {
for (i, s) in ss.iter_mut().enumerate() {
*s += src_pixel.0[i] as i32 * (k as i32);
}
}
for (i, s) in ss.iter().copied().enumerate() {
dst_pixel.0[i] = unsafe { normalizer.clip(s) };
}
dst_pixel.0 = ss.map(|v| unsafe { normalizer.clip(v) });
}
}
}
+15 -19
View File
@@ -3,16 +3,15 @@ use std::intrinsics::transmute;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::U8x4;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::{simd_utils, ImageView, ImageViewMut};
// This code is based on C-implementation from Pillow-SIMD package for Python
// https://github.com/uploadcare/pillow-simd
#[inline]
pub(crate) fn horiz_convolution(
src_image: &ImageView<U8x4>,
dst_image: &mut ImageViewMut<U8x4>,
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
offset: u32,
coeffs: Coefficients,
) {
@@ -21,39 +20,36 @@ pub(crate) fn horiz_convolution(
macro_rules! call {
($imm8:expr) => {{
horiz_convolution_p::<$imm8>(src_image, dst_image, offset, normalizer);
horiz_convolution_p::<$imm8>(src_view, dst_view, offset, normalizer);
}};
}
constify_imm8!(precision, call);
}
fn horiz_convolution_p<const PRECISION: i32>(
src_image: &ImageView<U8x4>,
dst_image: &mut ImageViewMut<U8x4>,
src_view: &impl ImageView<Pixel = U8x4>,
dst_view: &mut impl ImageViewMut<Pixel = U8x4>,
offset: u32,
normalizer: optimisations::Normalizer16,
) {
let coefficients_chunks = normalizer.normalized_chunks();
let dst_height = dst_image.height().get();
let dst_height = dst_view.height();
let src_iter = src_image.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_image.iter_4_rows_mut();
let src_iter = src_view.iter_4_rows(offset, dst_height + offset);
let dst_iter = dst_view.iter_4_rows_mut();
for (src_rows, dst_rows) in src_iter.zip(dst_iter) {
unsafe {
horiz_convolution_four_rows::<PRECISION>(src_rows, dst_rows, &coefficients_chunks);
}
}
let mut yy = dst_height - dst_height % 4;
while yy < dst_height {
let yy = dst_height - dst_height % 4;
let src_rows = src_view.iter_rows(yy + offset);
let dst_rows = dst_view.iter_rows_mut(yy);
for (src_row, dst_row) in src_rows.zip(dst_rows) {
unsafe {
horiz_convolution_one_row::<PRECISION>(
src_image.get_row(yy + offset).unwrap(),
dst_image.get_row_mut(yy).unwrap(),
&coefficients_chunks,
);
horiz_convolution_one_row::<PRECISION>(src_row, dst_row, &coefficients_chunks);
}
yy += 1;
}
}
@@ -66,7 +62,7 @@ fn horiz_convolution_p<const PRECISION: i32>(
#[target_feature(enable = "sse4.1")]
unsafe fn horiz_convolution_four_rows<const PRECISION: i32>(
src_rows: [&[U8x4]; 4],
dst_rows: [&mut &mut [U8x4]; 4],
dst_rows: [&mut [U8x4]; 4],
coefficients_chunks: &[optimisations::CoefficientsI16Chunk],
) {
let initial = _mm_set1_epi32(1 << (PRECISION - 1));
+11 -12
View File
@@ -1,39 +1,38 @@
use std::arch::x86_64::*;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::PixelExt;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::pixels::InnerPixel;
use crate::{simd_utils, ImageView, ImageViewMut};
pub(crate) fn vert_convolution<T>(
src_image: &ImageView<T>,
dst_image: &mut ImageViewMut<T>,
src_view: &impl ImageView<Pixel = T>,
dst_view: &mut impl ImageViewMut<Pixel = T>,
offset: u32,
coeffs: Coefficients,
) where
T: PixelExt<Component = u16>,
T: InnerPixel<Component = u16>,
{
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let src_x = offset as usize * T::count_of_components();
let dst_rows = dst_image.iter_rows_mut();
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, coeffs_chunk) in dst_rows.zip(coefficients_chunks) {
unsafe {
vert_convolution_into_one_row_u16(src_image, dst_row, src_x, coeffs_chunk, &normalizer);
vert_convolution_into_one_row_u16(src_view, dst_row, src_x, coeffs_chunk, &normalizer);
}
}
}
#[target_feature(enable = "avx2")]
pub(crate) unsafe fn vert_convolution_into_one_row_u16<T>(
src_img: &ImageView<T>,
src_view: &impl ImageView<Pixel = T>,
dst_row: &mut [T],
mut src_x: usize,
coeffs_chunk: optimisations::CoefficientsI32Chunk,
normalizer: &optimisations::Normalizer32,
) where
T: PixelExt<Component = u16>,
T: InnerPixel<Component = u16>,
{
let y_start = coeffs_chunk.start;
let coeffs = coeffs_chunk.values;
@@ -93,7 +92,7 @@ pub(crate) unsafe fn vert_convolution_into_one_row_u16<T>(
// 16 components / 4 per register = 4 registers
let mut sum = [initial; 4];
for (s_row, &coeff) in src_img.iter_rows(y_start).zip(coeffs) {
for (s_row, &coeff) in src_view.iter_rows(y_start).zip(coeffs) {
let components = T::components(s_row);
let coeff_i64x4 = _mm256_set1_epi64x(coeff as i64);
let source = simd_utils::loadu_si256(components, src_x);
@@ -124,7 +123,7 @@ pub(crate) unsafe fn vert_convolution_into_one_row_u16<T>(
let mut sum = [initial; 4];
let mut buf = [0u16; 16];
for (s_row, &coeff) in src_img.iter_rows(y_start).zip(coeffs) {
for (s_row, &coeff) in src_view.iter_rows(y_start).zip(coeffs) {
let components = T::components(s_row);
for (i, &v) in components
.get_unchecked(src_x..)
+13 -14
View File
@@ -1,7 +1,6 @@
use crate::convolution::Coefficients;
use crate::pixels::PixelExt;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::pixels::InnerPixel;
use crate::{CpuExtensions, ImageView, ImageViewMut};
#[cfg(target_arch = "x86_64")]
pub(crate) mod avx2;
@@ -11,28 +10,28 @@ mod neon;
#[cfg(target_arch = "x86_64")]
pub(crate) mod sse4;
#[cfg(target_arch = "wasm32")]
pub(crate) mod wasm32;
pub mod wasm32;
pub(crate) fn vert_convolution_u16<T: PixelExt<Component = u16>>(
src_image: &ImageView<T>,
dst_image: &mut ImageViewMut<T>,
pub(crate) fn vert_convolution_u16<T: InnerPixel<Component = u16>>(
src_view: &impl ImageView<Pixel = T>,
dst_view: &mut impl ImageViewMut<Pixel = T>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
// Check safety conditions
debug_assert!(src_image.width().get() - offset >= dst_image.width().get());
debug_assert_eq!(coeffs.bounds.len(), dst_image.height().get() as usize);
debug_assert!(src_view.width() - offset >= dst_view.width());
debug_assert_eq!(coeffs.bounds.len(), dst_view.height() as usize);
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => avx2::vert_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Avx2 => avx2::vert_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => sse4::vert_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Sse4_1 => sse4::vert_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => neon::vert_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Neon => neon::vert_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => wasm32::vert_convolution(src_image, dst_image, offset, coeffs),
_ => native::vert_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Simd128 => wasm32::vert_convolution(src_view, dst_view, offset, coeffs),
_ => native::vert_convolution(src_view, dst_view, offset, coeffs),
}
}
+13 -21
View File
@@ -1,16 +1,16 @@
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::PixelExt;
use crate::pixels::InnerPixel;
use crate::utils::foreach_with_pre_reading;
use crate::{ImageView, ImageViewMut};
#[inline(always)]
pub(crate) fn vert_convolution<T>(
src_image: &ImageView<T>,
dst_image: &mut ImageViewMut<T>,
src_view: &impl ImageView<Pixel = T>,
dst_view: &mut impl ImageViewMut<Pixel = T>,
offset: u32,
coeffs: Coefficients,
) where
T: PixelExt<Component = u16>,
T: InnerPixel<Component = u16>,
{
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
@@ -18,7 +18,7 @@ pub(crate) fn vert_convolution<T>(
let initial: i64 = 1 << (precision - 1);
let src_x_initial = offset as usize * T::count_of_components();
let dst_rows = dst_image.iter_rows_mut();
let dst_rows = dst_view.iter_rows_mut(0);
let coeffs_chunks_iter = coefficients_chunks.into_iter();
for (coeffs_chunk, dst_row) in coeffs_chunks_iter.zip(dst_rows) {
let first_y_src = coeffs_chunk.start;
@@ -28,7 +28,7 @@ pub(crate) fn vert_convolution<T>(
let (_, dst_chunks, tail) = unsafe { dst_components.align_to_mut::<[u16; 16]>() };
x_src = convolution_by_chunks(
src_image,
src_view,
&normalizer,
initial,
dst_chunks,
@@ -38,22 +38,14 @@ pub(crate) fn vert_convolution<T>(
);
if !tail.is_empty() {
convolution_by_u16(
src_image,
&normalizer,
initial,
tail,
x_src,
first_y_src,
ks,
);
convolution_by_u16(src_view, &normalizer, initial, tail, x_src, first_y_src, ks);
}
}
}
#[inline(always)]
pub(crate) fn convolution_by_u16<T: PixelExt<Component = u16>>(
src_image: &ImageView<T>,
pub(crate) fn convolution_by_u16<T: InnerPixel<Component = u16>>(
src_view: &impl ImageView<Pixel = T>,
normalizer: &optimisations::Normalizer32,
initial: i64,
dst_components: &mut [u16],
@@ -63,7 +55,7 @@ pub(crate) fn convolution_by_u16<T: PixelExt<Component = u16>>(
) -> usize {
for dst_component in dst_components.iter_mut() {
let mut ss = initial;
let src_rows = src_image.iter_rows(first_y_src);
let src_rows = src_view.iter_rows(first_y_src);
for (&k, src_row) in ks.iter().zip(src_rows) {
// SAFETY: Alignment of src_row is greater or equal than alignment u16
// because one component of pixel type T is u16.
@@ -79,7 +71,7 @@ pub(crate) fn convolution_by_u16<T: PixelExt<Component = u16>>(
#[inline(always)]
fn convolution_by_chunks<T, const CHUNK_SIZE: usize>(
src_image: &ImageView<T>,
src_view: &impl ImageView<Pixel = T>,
normalizer: &optimisations::Normalizer32,
initial: i64,
dst_chunks: &mut [[u16; CHUNK_SIZE]],
@@ -88,11 +80,11 @@ fn convolution_by_chunks<T, const CHUNK_SIZE: usize>(
ks: &[i32],
) -> usize
where
T: PixelExt<Component = u16>,
T: InnerPixel<Component = u16>,
{
for dst_chunk in dst_chunks {
let mut ss = [initial; CHUNK_SIZE];
let src_rows = src_image.iter_rows(first_y_src);
let src_rows = src_view.iter_rows(first_y_src);
foreach_with_pre_reading(
ks.iter().zip(src_rows),
+43 -39
View File
@@ -3,31 +3,32 @@ use std::arch::x86_64::*;
use crate::convolution::optimisations::CoefficientsI32Chunk;
use crate::convolution::vertical_u16::native::convolution_by_u16;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::PixelExt;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::pixels::InnerPixel;
use crate::{simd_utils, ImageView, ImageViewMut};
pub(crate) fn vert_convolution<T: PixelExt<Component = u16>>(
src_image: &ImageView<T>,
dst_image: &mut ImageViewMut<T>,
pub(crate) fn vert_convolution<T>(
src_view: &impl ImageView<Pixel = T>,
dst_view: &mut impl ImageViewMut<Pixel = T>,
offset: u32,
coeffs: Coefficients,
) {
) where
T: InnerPixel<Component = u16>,
{
let normalizer = optimisations::Normalizer32::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
let src_x = offset as usize * T::count_of_components();
let dst_rows = dst_image.iter_rows_mut();
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, coeffs_chunk) in dst_rows.zip(coefficients_chunks) {
unsafe {
vert_convolution_into_one_row_u16(src_image, dst_row, src_x, coeffs_chunk, &normalizer);
vert_convolution_into_one_row_u16(src_view, dst_row, src_x, coeffs_chunk, &normalizer);
}
}
}
#[target_feature(enable = "sse4.1")]
unsafe fn vert_convolution_into_one_row_u16<T: PixelExt<Component = u16>>(
src_img: &ImageView<T>,
unsafe fn vert_convolution_into_one_row_u16<T: InnerPixel<Component = u16>>(
src_view: &impl ImageView<Pixel = T>,
dst_row: &mut [T],
mut src_x: usize,
coeffs_chunk: CoefficientsI32Chunk,
@@ -35,7 +36,7 @@ unsafe fn vert_convolution_into_one_row_u16<T: PixelExt<Component = u16>>(
) {
let y_start = coeffs_chunk.start;
let coeffs = coeffs_chunk.values;
let max_y = y_start + coeffs.len() as u32;
let max_rows = coeffs.len() as u32;
let mut dst_u16 = T::components_mut(dst_row);
/*
@@ -77,7 +78,7 @@ unsafe fn vert_convolution_into_one_row_u16<T: PixelExt<Component = u16>>(
let coeffs_2 = coeffs.chunks_exact(2);
let coeffs_reminder = coeffs_2.remainder();
for (src_rows, two_coeffs) in src_img.iter_2_rows(y_start, max_y).zip(coeffs_2) {
for (src_rows, two_coeffs) in src_view.iter_2_rows(y_start, max_rows).zip(coeffs_2) {
let src_rows = src_rows.map(|row| T::components(row));
for r in 0..2 {
@@ -94,15 +95,16 @@ unsafe fn vert_convolution_into_one_row_u16<T: PixelExt<Component = u16>>(
}
if let Some(&k) = coeffs_reminder.first() {
let s_row = src_img.get_row(y_start + y).unwrap();
let components = T::components(s_row);
let coeff_i64x2 = _mm_set1_epi64x(k as i64);
if let Some(s_row) = src_view.iter_rows(y_start + y).next() {
let components = T::components(s_row);
let coeff_i64x2 = _mm_set1_epi64x(k as i64);
for x in 0..2 {
let source = simd_utils::loadu_si128(components, src_x + x * 8);
for i in 0..4 {
let c_i64x2 = _mm_shuffle_epi8(source, c_shuffles[i]);
sums[i][x] = _mm_add_epi64(sums[i][x], _mm_mul_epi32(c_i64x2, coeff_i64x2));
for x in 0..2 {
let source = simd_utils::loadu_si128(components, src_x + x * 8);
for i in 0..4 {
let c_i64x2 = _mm_shuffle_epi8(source, c_shuffles[i]);
sums[i][x] = _mm_add_epi64(sums[i][x], _mm_mul_epi32(c_i64x2, coeff_i64x2));
}
}
}
}
@@ -130,7 +132,7 @@ unsafe fn vert_convolution_into_one_row_u16<T: PixelExt<Component = u16>>(
let coeffs_2 = coeffs.chunks_exact(2);
let coeffs_reminder = coeffs_2.remainder();
for (src_rows, two_coeffs) in src_img.iter_2_rows(y_start, max_y).zip(coeffs_2) {
for (src_rows, two_coeffs) in src_view.iter_2_rows(y_start, max_rows).zip(coeffs_2) {
let src_rows = src_rows.map(|row| T::components(row));
let coeffs_i64 = [
_mm_set1_epi64x(two_coeffs[0] as i64),
@@ -148,13 +150,14 @@ unsafe fn vert_convolution_into_one_row_u16<T: PixelExt<Component = u16>>(
}
if let Some(&k) = coeffs_reminder.first() {
let s_row = src_img.get_row(y_start + y).unwrap();
let components = T::components(s_row);
let coeff_i64x2 = _mm_set1_epi64x(k as i64);
let source = simd_utils::loadu_si128(components, src_x);
for i in 0..4 {
let c_i64x2 = _mm_shuffle_epi8(source, c_shuffles[i]);
sums[i] = _mm_add_epi64(sums[i], _mm_mul_epi32(c_i64x2, coeff_i64x2));
if let Some(s_row) = src_view.iter_rows(y_start + y).next() {
let components = T::components(s_row);
let coeff_i64x2 = _mm_set1_epi64x(k as i64);
let source = simd_utils::loadu_si128(components, src_x);
for i in 0..4 {
let c_i64x2 = _mm_shuffle_epi8(source, c_shuffles[i]);
sums[i] = _mm_add_epi64(sums[i], _mm_mul_epi32(c_i64x2, coeff_i64x2));
}
}
}
@@ -183,7 +186,7 @@ unsafe fn vert_convolution_into_one_row_u16<T: PixelExt<Component = u16>>(
let coeffs_2 = coeffs.chunks_exact(2);
let coeffs_reminder = coeffs_2.remainder();
for (src_rows, two_coeffs) in src_img.iter_2_rows(y_start, max_y).zip(coeffs_2) {
for (src_rows, two_coeffs) in src_view.iter_2_rows(y_start, max_rows).zip(coeffs_2) {
let src_rows = src_rows.map(|row| T::components(row));
let coeffs_i64 = [
_mm_set1_epi64x(two_coeffs[0] as i64),
@@ -200,15 +203,16 @@ unsafe fn vert_convolution_into_one_row_u16<T: PixelExt<Component = u16>>(
}
if let Some(&k) = coeffs_reminder.first() {
let s_row = src_img.get_row(y_start + y).unwrap();
let components = T::components(s_row);
let coeff_i64x2 = _mm_set1_epi64x(k as i64);
if let Some(s_row) = src_view.iter_rows(y_start + y).next() {
let components = T::components(s_row);
let coeff_i64x2 = _mm_set1_epi64x(k as i64);
let comp_x4 = components.get_unchecked(src_x..src_x + 4);
let c_i64x2 = _mm_set_epi64x(comp_x4[1] as i64, comp_x4[0] as i64);
c01 = _mm_add_epi64(c01, _mm_mul_epi32(c_i64x2, coeff_i64x2));
let c_i64x2 = _mm_set_epi64x(comp_x4[3] as i64, comp_x4[2] as i64);
c23 = _mm_add_epi64(c23, _mm_mul_epi32(c_i64x2, coeff_i64x2));
let comp_x4 = components.get_unchecked(src_x..src_x + 4);
let c_i64x2 = _mm_set_epi64x(comp_x4[1] as i64, comp_x4[0] as i64);
c01 = _mm_add_epi64(c01, _mm_mul_epi32(c_i64x2, coeff_i64x2));
let c_i64x2 = _mm_set_epi64x(comp_x4[3] as i64, comp_x4[2] as i64);
c23 = _mm_add_epi64(c23, _mm_mul_epi32(c_i64x2, coeff_i64x2));
}
}
let mut dst_ptr = dst_chunk.as_mut_ptr();
@@ -229,7 +233,7 @@ unsafe fn vert_convolution_into_one_row_u16<T: PixelExt<Component = u16>>(
if !dst_u16.is_empty() {
let initial = 1 << (precision - 1);
convolution_by_u16(
src_img, normalizer, initial, dst_u16, src_x, y_start, coeffs,
src_view, normalizer, initial, dst_u16, src_x, y_start, coeffs,
);
}
}
+52 -49
View File
@@ -2,46 +2,46 @@ use std::arch::x86_64::*;
use crate::convolution::vertical_u8::native;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::PixelExt;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::image_view::ImageViewMut;
use crate::pixels::InnerPixel;
use crate::{simd_utils, ImageView};
#[inline]
pub(crate) fn vert_convolution<T>(
src_image: &ImageView<T>,
dst_image: &mut ImageViewMut<T>,
src_view: &impl ImageView<Pixel = T>,
dst_view: &mut impl ImageViewMut<Pixel = T>,
offset: u32,
coeffs: Coefficients,
) where
T: PixelExt<Component = u8>,
T: InnerPixel<Component = u8>,
{
let normalizer = optimisations::Normalizer16::new(coeffs);
let precision = normalizer.precision();
macro_rules! call {
($imm8:expr) => {{
vert_convolution_p::<T, $imm8>(src_image, dst_image, offset, normalizer);
vert_convolution_p::<T, $imm8>(src_view, dst_view, offset, normalizer);
}};
}
constify_imm8!(precision, call);
}
fn vert_convolution_p<T, const PRECISION: i32>(
src_image: &ImageView<T>,
dst_image: &mut ImageViewMut<T>,
src_view: &impl ImageView<Pixel = T>,
dst_view: &mut impl ImageViewMut<Pixel = T>,
offset: u32,
normalizer: optimisations::Normalizer16,
) where
T: PixelExt<Component = u8>,
T: InnerPixel<Component = u8>,
{
let coefficients_chunks = normalizer.normalized_chunks();
let src_x = offset as usize * T::count_of_components();
let dst_rows = dst_image.iter_rows_mut();
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, coeffs_chunk) in dst_rows.zip(coefficients_chunks) {
unsafe {
vert_convolution_into_one_row::<T, PRECISION>(
src_image,
src_view,
dst_row,
src_x,
coeffs_chunk,
@@ -54,17 +54,17 @@ fn vert_convolution_p<T, const PRECISION: i32>(
#[inline]
#[target_feature(enable = "avx2")]
unsafe fn vert_convolution_into_one_row<T, const PRECISION: i32>(
src_img: &ImageView<T>,
src_view: &impl ImageView<Pixel = T>,
dst_row: &mut [T],
mut src_x: usize,
coeffs_chunk: optimisations::CoefficientsI16Chunk,
normalizer: &optimisations::Normalizer16,
) where
T: PixelExt<Component = u8>,
T: InnerPixel<Component = u8>,
{
let y_start = coeffs_chunk.start;
let coeffs = coeffs_chunk.values;
let max_y = y_start + coeffs.len() as u32;
let max_rows = coeffs.len() as u32;
let initial = _mm_set1_epi32(1 << (PRECISION as u8 - 1));
let initial_256 = _mm256_set1_epi32(1 << (PRECISION as u8 - 1));
@@ -81,7 +81,7 @@ unsafe fn vert_convolution_into_one_row<T, const PRECISION: i32>(
let mut y: u32 = 0;
for src_rows in src_img.iter_2_rows(y_start, max_y) {
for src_rows in src_view.iter_2_rows(y_start, max_rows) {
let components1 = T::components(src_rows[0]);
let components2 = T::components(src_rows[1]);
@@ -107,24 +107,25 @@ unsafe fn vert_convolution_into_one_row<T, const PRECISION: i32>(
}
if let Some(&k) = coeffs.get(y as usize) {
let s_row = src_img.get_row(y_start + y).unwrap();
let components = T::components(s_row);
let mmk = _mm256_set1_epi32(k as i32);
if let Some(s_row) = src_view.iter_rows(y_start + y).next() {
let components = T::components(s_row);
let mmk = _mm256_set1_epi32(k as i32);
let source1 = simd_utils::loadu_si256(components, src_x); // top line
let source2 = _mm256_setzero_si256(); // bottom line is empty
let source1 = simd_utils::loadu_si256(components, src_x); // top line
let source2 = _mm256_setzero_si256(); // bottom line is empty
let source = _mm256_unpacklo_epi8(source1, source2);
let pix = _mm256_unpacklo_epi8(source, _mm256_setzero_si256());
sss0 = _mm256_add_epi32(sss0, _mm256_madd_epi16(pix, mmk));
let pix = _mm256_unpackhi_epi8(source, _mm256_setzero_si256());
sss1 = _mm256_add_epi32(sss1, _mm256_madd_epi16(pix, mmk));
let source = _mm256_unpacklo_epi8(source1, source2);
let pix = _mm256_unpacklo_epi8(source, _mm256_setzero_si256());
sss0 = _mm256_add_epi32(sss0, _mm256_madd_epi16(pix, mmk));
let pix = _mm256_unpackhi_epi8(source, _mm256_setzero_si256());
sss1 = _mm256_add_epi32(sss1, _mm256_madd_epi16(pix, mmk));
let source = _mm256_unpackhi_epi8(source1, _mm256_setzero_si256());
let pix = _mm256_unpacklo_epi8(source, _mm256_setzero_si256());
sss2 = _mm256_add_epi32(sss2, _mm256_madd_epi16(pix, mmk));
let pix = _mm256_unpackhi_epi8(source, _mm256_setzero_si256());
sss3 = _mm256_add_epi32(sss3, _mm256_madd_epi16(pix, mmk));
let source = _mm256_unpackhi_epi8(source1, _mm256_setzero_si256());
let pix = _mm256_unpacklo_epi8(source, _mm256_setzero_si256());
sss2 = _mm256_add_epi32(sss2, _mm256_madd_epi16(pix, mmk));
let pix = _mm256_unpackhi_epi8(source, _mm256_setzero_si256());
sss3 = _mm256_add_epi32(sss3, _mm256_madd_epi16(pix, mmk));
}
}
sss0 = _mm256_srai_epi32::<PRECISION>(sss0);
@@ -150,7 +151,7 @@ unsafe fn vert_convolution_into_one_row<T, const PRECISION: i32>(
let mut sss1 = initial; // right row
let mut y: u32 = 0;
for src_rows in src_img.iter_2_rows(y_start, max_y) {
for src_rows in src_view.iter_2_rows(y_start, max_rows) {
let components1 = T::components(src_rows[0]);
let components2 = T::components(src_rows[1]);
// Load two coefficients at once
@@ -169,18 +170,19 @@ unsafe fn vert_convolution_into_one_row<T, const PRECISION: i32>(
}
if let Some(&k) = coeffs.get(y as usize) {
let s_row = src_img.get_row(y_start + y).unwrap();
let components = T::components(s_row);
let mmk = _mm_set1_epi32(k as i32);
if let Some(s_row) = src_view.iter_rows(y_start + y).next() {
let components = T::components(s_row);
let mmk = _mm_set1_epi32(k as i32);
let source1 = simd_utils::loadl_epi64(components, src_x); // top line
let source2 = _mm_setzero_si128(); // bottom line is empty
let source1 = simd_utils::loadl_epi64(components, src_x); // top line
let source2 = _mm_setzero_si128(); // bottom line is empty
let source = _mm_unpacklo_epi8(source1, source2);
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss0 = _mm_add_epi32(sss0, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss1 = _mm_add_epi32(sss1, _mm_madd_epi16(pix, mmk));
let source = _mm_unpacklo_epi8(source1, source2);
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss0 = _mm_add_epi32(sss0, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss1 = _mm_add_epi32(sss1, _mm_madd_epi16(pix, mmk));
}
}
sss0 = _mm_srai_epi32::<PRECISION>(sss0);
@@ -200,7 +202,7 @@ unsafe fn vert_convolution_into_one_row<T, const PRECISION: i32>(
if let Some(dst_chunk) = dst_chunks_4.next() {
let mut sss = initial;
let mut y: u32 = 0;
for src_rows in src_img.iter_2_rows(y_start, max_y) {
for src_rows in src_view.iter_2_rows(y_start, max_rows) {
let components1 = T::components(src_rows[0]);
let components2 = T::components(src_rows[1]);
// Load two coefficients at once
@@ -217,11 +219,12 @@ unsafe fn vert_convolution_into_one_row<T, const PRECISION: i32>(
}
if let Some(&k) = coeffs.get(y as usize) {
let s_row = src_img.get_row(y_start + y).unwrap();
let components = T::components(s_row);
let pix = simd_utils::mm_cvtepu8_epi32_from_u8(components, src_x);
let mmk = _mm_set1_epi32(k as i32);
sss = _mm_add_epi32(sss, _mm_madd_epi16(pix, mmk));
if let Some(s_row) = src_view.iter_rows(y_start + y).next() {
let components = T::components(s_row);
let pix = simd_utils::mm_cvtepu8_epi32_from_u8(components, src_x);
let mmk = _mm_set1_epi32(k as i32);
sss = _mm_add_epi32(sss, _mm_madd_epi16(pix, mmk));
}
}
sss = _mm_srai_epi32::<PRECISION>(sss);
@@ -236,7 +239,7 @@ unsafe fn vert_convolution_into_one_row<T, const PRECISION: i32>(
dst_u8 = dst_chunks_4.into_remainder();
if !dst_u8.is_empty() {
native::convolution_by_u8(
src_img,
src_view,
normalizer,
1 << (PRECISION as u8 - 1),
dst_u8,
+12 -13
View File
@@ -1,7 +1,6 @@
use crate::convolution::Coefficients;
use crate::pixels::PixelExt;
use crate::CpuExtensions;
use crate::{ImageView, ImageViewMut};
use crate::pixels::InnerPixel;
use crate::{CpuExtensions, ImageView, ImageViewMut};
#[cfg(target_arch = "x86_64")]
pub(crate) mod avx2;
@@ -13,26 +12,26 @@ pub(crate) mod sse4;
#[cfg(target_arch = "wasm32")]
pub(crate) mod wasm32;
pub(crate) fn vert_convolution_u8<T: PixelExt<Component = u8>>(
src_image: &ImageView<T>,
dst_image: &mut ImageViewMut<T>,
pub(crate) fn vert_convolution_u8<T: InnerPixel<Component = u8>>(
src_view: &impl ImageView<Pixel = T>,
dst_view: &mut impl ImageViewMut<Pixel = T>,
offset: u32,
coeffs: Coefficients,
cpu_extensions: CpuExtensions,
) {
// Check safety conditions
debug_assert!(src_image.width().get() - offset >= dst_image.width().get());
debug_assert_eq!(coeffs.bounds.len(), dst_image.height().get() as usize);
debug_assert!(src_view.width() - offset >= dst_view.width());
debug_assert_eq!(coeffs.bounds.len(), dst_view.height() as usize);
match cpu_extensions {
#[cfg(target_arch = "x86_64")]
CpuExtensions::Avx2 => avx2::vert_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Avx2 => avx2::vert_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "x86_64")]
CpuExtensions::Sse4_1 => sse4::vert_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Sse4_1 => sse4::vert_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "aarch64")]
CpuExtensions::Neon => neon::vert_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Neon => neon::vert_convolution(src_view, dst_view, offset, coeffs),
#[cfg(target_arch = "wasm32")]
CpuExtensions::Simd128 => wasm32::vert_convolution(src_image, dst_image, offset, coeffs),
_ => native::vert_convolution(src_image, dst_image, offset, coeffs),
CpuExtensions::Simd128 => wasm32::vert_convolution(src_view, dst_view, offset, coeffs),
_ => native::vert_convolution(src_view, dst_view, offset, coeffs),
}
}
+10 -10
View File
@@ -1,16 +1,16 @@
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::PixelExt;
use crate::image_view::{ImageView, ImageViewMut};
use crate::pixels::InnerPixel;
use crate::utils::foreach_with_pre_reading;
use crate::{ImageView, ImageViewMut};
#[inline(always)]
pub(crate) fn vert_convolution<T>(
src_image: &ImageView<T>,
dst_image: &mut ImageViewMut<T>,
src_image: &impl ImageView<Pixel = T>,
dst_image: &mut impl ImageViewMut<Pixel = T>,
offset: u32,
coeffs: Coefficients,
) where
T: PixelExt<Component = u8>,
T: InnerPixel<Component = u8>,
{
let normalizer = optimisations::Normalizer16::new(coeffs);
let coefficients_chunks = normalizer.normalized_chunks();
@@ -18,7 +18,7 @@ pub(crate) fn vert_convolution<T>(
let initial = 1 << (precision - 1);
let src_x_initial = offset as usize * T::count_of_components();
let dst_rows = dst_image.iter_rows_mut();
let dst_rows = dst_image.iter_rows_mut(0);
let coeffs_chunks_iter = coefficients_chunks.into_iter();
for (coeffs_chunk, dst_row) in coeffs_chunks_iter.zip(dst_rows) {
let first_y_src = coeffs_chunk.start;
@@ -95,7 +95,7 @@ pub(crate) fn vert_convolution<T>(
#[inline(always)]
pub(crate) fn convolution_by_u8<T>(
src_image: &ImageView<T>,
src_image: &impl ImageView<Pixel = T>,
normalizer: &optimisations::Normalizer16,
initial: i32,
dst_components: &mut [u8],
@@ -104,7 +104,7 @@ pub(crate) fn convolution_by_u8<T>(
ks: &[i16],
) -> usize
where
T: PixelExt<Component = u8>,
T: InnerPixel<Component = u8>,
{
for dst_component in dst_components {
let mut ss = initial;
@@ -122,7 +122,7 @@ where
#[inline(always)]
fn convolution_by_chunks<T, const CHUNK_SIZE: usize>(
src_image: &ImageView<T>,
src_image: &impl ImageView<Pixel = T>,
normalizer: &optimisations::Normalizer16,
initial: i32,
dst_chunks: &mut [[u8; CHUNK_SIZE]],
@@ -131,7 +131,7 @@ fn convolution_by_chunks<T, const CHUNK_SIZE: usize>(
ks: &[i16],
) -> usize
where
T: PixelExt<Component = u8>,
T: InnerPixel<Component = u8>,
{
for dst_chunk in dst_chunks {
let mut ss = [initial; CHUNK_SIZE];
+63 -59
View File
@@ -2,46 +2,45 @@ use std::arch::x86_64::*;
use crate::convolution::vertical_u8::native;
use crate::convolution::{optimisations, Coefficients};
use crate::pixels::PixelExt;
use crate::simd_utils;
use crate::{ImageView, ImageViewMut};
use crate::pixels::InnerPixel;
use crate::{simd_utils, ImageView, ImageViewMut};
#[inline]
pub(crate) fn vert_convolution<T>(
src_image: &ImageView<T>,
dst_image: &mut ImageViewMut<T>,
src_view: &impl ImageView<Pixel = T>,
dst_view: &mut impl ImageViewMut<Pixel = T>,
offset: u32,
coeffs: Coefficients,
) where
T: PixelExt<Component = u8>,
T: InnerPixel<Component = u8>,
{
let normalizer = optimisations::Normalizer16::new(coeffs);
let precision = normalizer.precision();
macro_rules! call {
($imm8:expr) => {{
vert_convolution_p::<T, $imm8>(src_image, dst_image, offset, normalizer);
vert_convolution_p::<T, $imm8>(src_view, dst_view, offset, normalizer);
}};
}
constify_imm8!(precision, call);
}
fn vert_convolution_p<T, const PRECISION: i32>(
src_image: &ImageView<T>,
dst_image: &mut ImageViewMut<T>,
src_view: &impl ImageView<Pixel = T>,
dst_view: &mut impl ImageViewMut<Pixel = T>,
offset: u32,
normalizer: optimisations::Normalizer16,
) where
T: PixelExt<Component = u8>,
T: InnerPixel<Component = u8>,
{
let coefficients_chunks = normalizer.normalized_chunks();
let src_x = offset as usize * T::count_of_components();
let dst_rows = dst_image.iter_rows_mut();
let dst_rows = dst_view.iter_rows_mut(0);
for (dst_row, coeffs_chunk) in dst_rows.zip(coefficients_chunks) {
unsafe {
vert_convolution_into_one_row::<T, PRECISION>(
src_image,
src_view,
dst_row,
src_x,
coeffs_chunk,
@@ -52,16 +51,18 @@ fn vert_convolution_p<T, const PRECISION: i32>(
}
#[target_feature(enable = "sse4.1")]
unsafe fn vert_convolution_into_one_row<T: PixelExt<Component = u8>, const PRECISION: i32>(
src_img: &ImageView<T>,
unsafe fn vert_convolution_into_one_row<T, const PRECISION: i32>(
src_view: &impl ImageView<Pixel = T>,
dst_row: &mut [T],
mut src_x: usize,
coeffs_chunk: optimisations::CoefficientsI16Chunk,
normalizer: &optimisations::Normalizer16,
) {
) where
T: InnerPixel<Component = u8>,
{
let y_start = coeffs_chunk.start;
let coeffs = coeffs_chunk.values;
let max_y = y_start + coeffs.len() as u32;
let max_rows = coeffs.len() as u32;
let mut dst_u8 = T::components_mut(dst_row);
let initial = _mm_set1_epi32(1 << (PRECISION - 1));
@@ -79,7 +80,7 @@ unsafe fn vert_convolution_into_one_row<T: PixelExt<Component = u8>, const PRECI
let mut y: u32 = 0;
for src_rows in src_img.iter_2_rows(y_start, max_y) {
for src_rows in src_view.iter_2_rows(y_start, max_rows) {
let components1 = T::components(src_rows[0]);
let components2 = T::components(src_rows[1]);
@@ -120,37 +121,38 @@ unsafe fn vert_convolution_into_one_row<T: PixelExt<Component = u8>, const PRECI
}
if let Some(&k) = coeffs.get(y as usize) {
let s_row = src_img.get_row(y_start + y).unwrap();
let components = T::components(s_row);
let mmk = _mm_set1_epi32(k as i32);
if let Some(s_row) = src_view.iter_rows(y_start + y).next() {
let components = T::components(s_row);
let mmk = _mm_set1_epi32(k as i32);
let source1 = simd_utils::loadu_si128(components, src_x); // top line
let source1 = simd_utils::loadu_si128(components, src_x); // top line
let source = _mm_unpacklo_epi8(source1, _mm_setzero_si128());
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss0 = _mm_add_epi32(sss0, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss1 = _mm_add_epi32(sss1, _mm_madd_epi16(pix, mmk));
let source = _mm_unpacklo_epi8(source1, _mm_setzero_si128());
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss0 = _mm_add_epi32(sss0, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss1 = _mm_add_epi32(sss1, _mm_madd_epi16(pix, mmk));
let source = _mm_unpackhi_epi8(source1, _mm_setzero_si128());
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss2 = _mm_add_epi32(sss2, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss3 = _mm_add_epi32(sss3, _mm_madd_epi16(pix, mmk));
let source = _mm_unpackhi_epi8(source1, _mm_setzero_si128());
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss2 = _mm_add_epi32(sss2, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss3 = _mm_add_epi32(sss3, _mm_madd_epi16(pix, mmk));
let source1 = simd_utils::loadu_si128(components, src_x + 16); // top line
let source1 = simd_utils::loadu_si128(components, src_x + 16); // top line
let source = _mm_unpacklo_epi8(source1, _mm_setzero_si128());
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss4 = _mm_add_epi32(sss4, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss5 = _mm_add_epi32(sss5, _mm_madd_epi16(pix, mmk));
let source = _mm_unpacklo_epi8(source1, _mm_setzero_si128());
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss4 = _mm_add_epi32(sss4, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss5 = _mm_add_epi32(sss5, _mm_madd_epi16(pix, mmk));
let source = _mm_unpackhi_epi8(source1, _mm_setzero_si128());
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss6 = _mm_add_epi32(sss6, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss7 = _mm_add_epi32(sss7, _mm_madd_epi16(pix, mmk));
let source = _mm_unpackhi_epi8(source1, _mm_setzero_si128());
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss6 = _mm_add_epi32(sss6, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss7 = _mm_add_epi32(sss7, _mm_madd_epi16(pix, mmk));
}
}
sss0 = _mm_srai_epi32::<PRECISION>(sss0);
@@ -183,7 +185,7 @@ unsafe fn vert_convolution_into_one_row<T: PixelExt<Component = u8>, const PRECI
let mut sss1 = initial; // right row
let mut y: u32 = 0;
for src_rows in src_img.iter_2_rows(y_start, max_y) {
for src_rows in src_view.iter_2_rows(y_start, max_rows) {
let components1 = T::components(src_rows[0]);
let components2 = T::components(src_rows[1]);
// Load two coefficients at once
@@ -202,17 +204,18 @@ unsafe fn vert_convolution_into_one_row<T: PixelExt<Component = u8>, const PRECI
}
if let Some(&k) = coeffs.get(y as usize) {
let s_row = src_img.get_row(y_start + y).unwrap();
let components = T::components(s_row);
let mmk = _mm_set1_epi32(k as i32);
if let Some(s_row) = src_view.iter_rows(y_start + y).next() {
let components = T::components(s_row);
let mmk = _mm_set1_epi32(k as i32);
let source1 = simd_utils::loadl_epi64(components, src_x); // top line
let source1 = simd_utils::loadl_epi64(components, src_x); // top line
let source = _mm_unpacklo_epi8(source1, _mm_setzero_si128());
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss0 = _mm_add_epi32(sss0, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss1 = _mm_add_epi32(sss1, _mm_madd_epi16(pix, mmk));
let source = _mm_unpacklo_epi8(source1, _mm_setzero_si128());
let pix = _mm_unpacklo_epi8(source, _mm_setzero_si128());
sss0 = _mm_add_epi32(sss0, _mm_madd_epi16(pix, mmk));
let pix = _mm_unpackhi_epi8(source, _mm_setzero_si128());
sss1 = _mm_add_epi32(sss1, _mm_madd_epi16(pix, mmk));
}
}
sss0 = _mm_srai_epi32::<PRECISION>(sss0);
@@ -232,7 +235,7 @@ unsafe fn vert_convolution_into_one_row<T: PixelExt<Component = u8>, const PRECI
let mut sss = initial;
let mut y: u32 = 0;
for src_rows in src_img.iter_2_rows(y_start, max_y) {
for src_rows in src_view.iter_2_rows(y_start, max_rows) {
let components1 = T::components(src_rows[0]);
let components2 = T::components(src_rows[1]);
// Load two coefficients at once
@@ -249,11 +252,12 @@ unsafe fn vert_convolution_into_one_row<T: PixelExt<Component = u8>, const PRECI
}
if let Some(&k) = coeffs.get(y as usize) {
let s_row = src_img.get_row(y_start + y).unwrap();
let components = T::components(s_row);
let pix = simd_utils::mm_cvtepu8_epi32_from_u8(components, src_x);
let mmk = _mm_set1_epi32(k as i32);
sss = _mm_add_epi32(sss, _mm_madd_epi16(pix, mmk));
if let Some(s_row) = src_view.iter_rows(y_start + y).next() {
let components = T::components(s_row);
let pix = simd_utils::mm_cvtepu8_epi32_from_u8(components, src_x);
let mmk = _mm_set1_epi32(k as i32);
sss = _mm_add_epi32(sss, _mm_madd_epi16(pix, mmk));
}
}
sss = _mm_srai_epi32::<PRECISION>(sss);
@@ -268,7 +272,7 @@ unsafe fn vert_convolution_into_one_row<T: PixelExt<Component = u8>, const PRECI
dst_u8 = dst_chunks_4.into_remainder();
if !dst_u8.is_empty() {
native::convolution_by_u8(
src_img,
src_view,
normalizer,
1 << (PRECISION - 1),
dst_u8,
+72
View File
@@ -0,0 +1,72 @@
/// SIMD extension of CPU.
/// Specific variants depend on target architecture.
/// Look at source code to see all available variants.
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
pub enum CpuExtensions {
None,
#[cfg(target_arch = "x86_64")]
/// SIMD extension of x86_64 architecture
Sse4_1,
#[cfg(target_arch = "x86_64")]
/// SIMD extension of x86_64 architecture
Avx2,
#[cfg(target_arch = "aarch64")]
/// SIMD extension of Arm64 architecture
Neon,
#[cfg(target_arch = "wasm32")]
/// SIMD extension of Wasm32 architecture
Simd128,
}
impl CpuExtensions {
/// Returns `true` if your CPU support the extension.
pub fn is_supported(&self) -> bool {
match self {
#[cfg(target_arch = "x86_64")]
Self::Avx2 => is_x86_feature_detected!("avx2"),
#[cfg(target_arch = "x86_64")]
Self::Sse4_1 => is_x86_feature_detected!("sse4.1"),
#[cfg(target_arch = "aarch64")]
Self::Neon => true,
#[cfg(target_arch = "wasm32")]
Self::Simd128 => true,
Self::None => true,
}
}
}
impl Default for CpuExtensions {
#[cfg(target_arch = "x86_64")]
fn default() -> Self {
if is_x86_feature_detected!("avx2") {
Self::Avx2
} else if is_x86_feature_detected!("sse4.1") {
Self::Sse4_1
} else {
Self::None
}
}
#[cfg(target_arch = "aarch64")]
fn default() -> Self {
use std::arch::is_aarch64_feature_detected;
if is_aarch64_feature_detected!("neon") {
Self::Neon
} else {
Self::None
}
}
#[cfg(target_arch = "wasm32")]
fn default() -> Self {
Self::Simd128
}
#[cfg(not(any(
target_arch = "x86_64",
target_arch = "aarch64",
target_arch = "wasm32"
)))]
fn default() -> Self {
Self::None
}
}
+141
View File
@@ -0,0 +1,141 @@
use crate::{CropBoxError, ImageView};
/// A crop box parameters.
#[derive(Debug, Clone, Copy)]
pub struct CropBox {
pub left: f64,
pub top: f64,
pub width: f64,
pub height: f64,
}
impl CropBox {
/// Get a crop box to resize the source image into the
/// aspect ratio of destination image without distortions.
///
/// `centering` used to control the cropping position. Use (0.5, 0.5) for
/// center cropping (e.g. if cropping the width, take 50% off
/// of the left side, and therefore 50% off the right side).
/// (0.0, 0.0) will crop from the top left corner (i.e. if
/// cropping the width, take all the crop off of the right
/// side, and if cropping the height, take all of it off the
/// bottom). (1.0, 0.0) will crop from the bottom left
/// corner, etc. (i.e. if cropping the width, take all the
/// crop off the left side, and if cropping the height take
/// none from the top, and therefore all off the bottom).
pub fn fit_src_into_dst_size(
src_width: u32,
src_height: u32,
dst_width: u32,
dst_height: u32,
centering: Option<(f64, f64)>,
) -> Self {
if src_width == 0 || src_height == 0 || dst_width == 0 || dst_height == 0 {
return Self {
left: 0.,
top: 0.,
width: src_width as _,
height: src_height as _,
};
}
// This function based on code of ImageOps.fit() from Pillow package.
// https://github.com/python-pillow/Pillow/blob/master/src/PIL/ImageOps.py
let centering = if let Some((x, y)) = centering {
(x.clamp(0.0, 1.0), y.clamp(0.0, 1.0))
} else {
(0.5, 0.5)
};
// calculate aspect ratios
let width = src_width as f64;
let height = src_height as f64;
let image_ratio = width / height;
let required_ration = dst_width as f64 / dst_height as f64;
let crop_width;
let crop_height;
// figure out if the sides or top/bottom will be cropped off
if (image_ratio - required_ration).abs() < f64::EPSILON {
// The image is already the needed ratio
crop_width = width;
crop_height = height;
} else if image_ratio >= required_ration {
// The image is wider than what's needed, crop the sides
crop_width = required_ration * height;
crop_height = height;
} else {
// The image is taller than what's needed, crop the top and bottom
crop_width = width;
crop_height = width / required_ration;
}
let crop_left = (width - crop_width) * centering.0;
let crop_top = (height - crop_height) * centering.1;
Self {
left: crop_left,
top: crop_top,
width: crop_width,
height: crop_height,
}
}
}
pub(crate) struct CroppedSrcImageView<'a, T: ImageView> {
image_view: &'a T,
crop_box: CropBox,
}
impl<'a, T: ImageView> CroppedSrcImageView<'a, T> {
pub fn new(image_view: &'a T) -> Self {
Self {
image_view,
crop_box: CropBox {
left: 0.0,
top: 0.0,
width: image_view.width() as _,
height: image_view.height() as _,
},
}
}
pub fn cropped(image_view: &'a T, crop_box: CropBox) -> Result<Self, CropBoxError> {
if crop_box.width <= 0. || crop_box.height <= 0. {
return Err(CropBoxError::WidthOrHeightLessOrEqualToZero);
}
let img_width = image_view.width() as _;
let img_height = image_view.height() as _;
if crop_box.left >= img_width || crop_box.top >= img_height {
return Err(CropBoxError::PositionIsOutOfImageBoundaries);
}
let right = crop_box.left + crop_box.width;
let bottom = crop_box.top + crop_box.height;
if right > img_width || bottom > img_height {
return Err(CropBoxError::SizeIsOutOfImageBoundaries);
}
Ok(Self {
image_view,
crop_box,
})
}
pub unsafe fn cropped_unchecked(image_view: &'a T, crop_box: CropBox) -> Self {
Self {
image_view,
crop_box,
}
}
#[inline]
pub fn image_view(&self) -> &T {
self.image_view
}
#[inline]
pub fn crop_box(&self) -> CropBox {
self.crop_box
}
}
-247
View File
@@ -1,247 +0,0 @@
use std::num::NonZeroU32;
use crate::image_view::change_type_of_pixel_components;
use crate::pixels::{U16x2, U16x3, U16x4, U8x2, U8x3, U8x4, F32, I32, U16, U8};
use crate::{CropBox, CropBoxError, ImageView, ImageViewMut, MappingError, PixelType};
/// An immutable view of image data used by resizer as source image.
#[derive(Debug, Clone)]
#[non_exhaustive]
pub enum DynamicImageView<'a> {
U8(ImageView<'a, U8>),
U8x2(ImageView<'a, U8x2>),
U8x3(ImageView<'a, U8x3>),
U8x4(ImageView<'a, U8x4>),
U16(ImageView<'a, U16>),
U16x2(ImageView<'a, U16x2>),
U16x3(ImageView<'a, U16x3>),
U16x4(ImageView<'a, U16x4>),
I32(ImageView<'a, I32>),
F32(ImageView<'a, F32>),
}
/// A mutable view of image data used by resizer as destination image.
#[derive(Debug)]
#[non_exhaustive]
pub enum DynamicImageViewMut<'a> {
U8(ImageViewMut<'a, U8>),
U8x2(ImageViewMut<'a, U8x2>),
U8x3(ImageViewMut<'a, U8x3>),
U8x4(ImageViewMut<'a, U8x4>),
U16(ImageViewMut<'a, U16>),
U16x2(ImageViewMut<'a, U16x2>),
U16x3(ImageViewMut<'a, U16x3>),
U16x4(ImageViewMut<'a, U16x4>),
I32(ImageViewMut<'a, I32>),
F32(ImageViewMut<'a, F32>),
}
macro_rules! dynamic_map(
($dyn_image: expr, $image: pat => $action: expr) => ({
use DynamicImageView::*;
match $dyn_image {
U8($image) => U8($action),
U8x2($image) => U8x2($action),
U8x3($image) => U8x3($action),
U8x4($image) => U8x4($action),
U16($image) => U16($action),
U16x2($image) => U16x2($action),
U16x3($image) => U16x3($action),
U16x4($image) => U16x4($action),
I32($image) => I32($action),
F32($image) => F32($action),
}
});
($dyn_image: expr, |$image: pat_param| $action: expr) => (
match $dyn_image {
DynamicImageView::U8($image) => $action,
DynamicImageView::U8x2($image) => $action,
DynamicImageView::U8x3($image) => $action,
DynamicImageView::U8x4($image) => $action,
DynamicImageView::U16($image) => $action,
DynamicImageView::U16x2($image) => $action,
DynamicImageView::U16x3($image) => $action,
DynamicImageView::U16x4($image) => $action,
DynamicImageView::I32($image) => $action,
DynamicImageView::F32($image) => $action,
}
);
);
macro_rules! dynamic_mut_map (
($dyn_image: expr, $image: pat => $action: expr) => ({
use DynamicImageViewMut::*;
match $dyn_image {
U8($image) => U8($action),
U8x2($image) => U8x2($action),
U8x3($image) => U8x3($action),
U8x4($image) => U8x4($action),
U16($image) => U16($action),
U16x2($image) => U16x2($action),
U16x3($image) => U16x3($action),
U16x4($image) => U16x4($action),
I32($image) => I32($action),
F32($image) => F32($action),
}
});
($dyn_image: expr, |$image: pat_param| $action: expr) => (
match $dyn_image {
DynamicImageViewMut::U8($image) => $action,
DynamicImageViewMut::U8x2($image) => $action,
DynamicImageViewMut::U8x3($image) => $action,
DynamicImageViewMut::U8x4($image) => $action,
DynamicImageViewMut::U16($image) => $action,
DynamicImageViewMut::U16x2($image) => $action,
DynamicImageViewMut::U16x3($image) => $action,
DynamicImageViewMut::U16x4($image) => $action,
DynamicImageViewMut::I32($image) => $action,
DynamicImageViewMut::F32($image) => $action,
}
);
);
impl<'a> DynamicImageView<'a> {
pub fn width(&self) -> NonZeroU32 {
dynamic_map!(self, |typed_image| typed_image.width())
}
pub fn height(&self) -> NonZeroU32 {
dynamic_map!(self, |typed_image| typed_image.height())
}
pub fn pixel_type(&self) -> PixelType {
dynamic_map!(self, |typed_image| typed_image.pixel_type())
}
pub fn crop_box(&self) -> CropBox {
dynamic_map!(self, |typed_image| typed_image.crop_box())
}
pub fn set_crop_box(&mut self, crop_box: CropBox) -> Result<(), CropBoxError> {
dynamic_map!(self, |typed_image| typed_image.set_crop_box(crop_box))
}
pub fn set_crop_box_to_fit_dst_size(
&mut self,
dst_width: NonZeroU32,
dst_height: NonZeroU32,
centering: Option<(f64, f64)>,
) {
dynamic_map!(self, |typed_image| typed_image
.set_crop_box_to_fit_dst_size(dst_width, dst_height, centering))
}
}
impl<'a> DynamicImageViewMut<'a> {
pub fn width(&self) -> NonZeroU32 {
dynamic_mut_map!(self, |typed_image| typed_image.width())
}
pub fn height(&self) -> NonZeroU32 {
dynamic_mut_map!(self, |typed_image| typed_image.height())
}
pub fn pixel_type(&self) -> PixelType {
dynamic_mut_map!(self, |typed_image| typed_image.pixel_type())
}
/// Create cropped version of the view.
pub fn crop(
self,
left: u32,
top: u32,
width: NonZeroU32,
height: NonZeroU32,
) -> Result<Self, CropBoxError> {
Ok(dynamic_mut_map!(
self,
typed_image => typed_image.crop(left, top, width, height)?
))
}
}
macro_rules! from_typed {
($pixel_type: ty, $enum: expr, $enum_mut: expr) => {
impl<'a> From<ImageView<'a, $pixel_type>> for DynamicImageView<'a> {
fn from(view: ImageView<'a, $pixel_type>) -> Self {
$enum(view)
}
}
impl<'a> From<ImageViewMut<'a, $pixel_type>> for DynamicImageViewMut<'a> {
fn from(view: ImageViewMut<'a, $pixel_type>) -> Self {
$enum_mut(view)
}
}
};
}
from_typed!(U8, DynamicImageView::U8, DynamicImageViewMut::U8);
from_typed!(U8x2, DynamicImageView::U8x2, DynamicImageViewMut::U8x2);
from_typed!(U8x3, DynamicImageView::U8x3, DynamicImageViewMut::U8x3);
from_typed!(U8x4, DynamicImageView::U8x4, DynamicImageViewMut::U8x4);
from_typed!(U16, DynamicImageView::U16, DynamicImageViewMut::U16);
from_typed!(U16x2, DynamicImageView::U16x2, DynamicImageViewMut::U16x2);
from_typed!(U16x3, DynamicImageView::U16x3, DynamicImageViewMut::U16x3);
from_typed!(U16x4, DynamicImageView::U16x4, DynamicImageViewMut::U16x4);
from_typed!(I32, DynamicImageView::I32, DynamicImageViewMut::I32);
from_typed!(F32, DynamicImageView::F32, DynamicImageViewMut::F32);
pub fn change_type_of_pixel_components_dyn(
src_image: &DynamicImageView,
dst_image: &mut DynamicImageViewMut,
) -> Result<(), MappingError> {
macro_rules! map {
($value:expr, $(($src_first:path, $src_second:path, $dst_first:path, $dst_second:path)),*) => {
match $value {
$(
($src_first(src), $dst_first(dst)) => {
change_type_of_pixel_components(src, dst)?;
}
($src_first(src), $dst_second(dst)) => {
change_type_of_pixel_components(src, dst)?;
}
($src_second(src), $dst_first(dst)) => {
change_type_of_pixel_components(src, dst)?;
}
($src_second(src), $dst_second(dst)) => {
change_type_of_pixel_components(src, dst)?;
}
)*
_ => return Err(MappingError::UnsupportedCombinationOfImageTypes),
}
}
}
use DynamicImageView as IV;
use DynamicImageViewMut as IVMut;
map!(
(src_image, dst_image),
(IV::U8, IV::U16, IVMut::U8, IVMut::U16),
(IV::U8x2, IV::U16x2, IVMut::U8x2, IVMut::U16x2),
(IV::U8x3, IV::U16x3, IVMut::U8x3, IVMut::U16x3),
(IV::U8x4, IV::U16x4, IVMut::U8x4, IVMut::U16x4)
);
Ok(())
}
impl<'a> From<DynamicImageViewMut<'a>> for DynamicImageView<'a> {
fn from(dyn_view: DynamicImageViewMut<'a>) -> Self {
use DynamicImageViewMut::*;
match dyn_view {
U8(typed_view) => DynamicImageView::U8(typed_view.into()),
U8x2(typed_view) => DynamicImageView::U8x2(typed_view.into()),
U8x3(typed_view) => DynamicImageView::U8x3(typed_view.into()),
U8x4(typed_view) => DynamicImageView::U8x4(typed_view.into()),
U16(typed_view) => DynamicImageView::U16(typed_view.into()),
U16x2(typed_view) => DynamicImageView::U16x2(typed_view.into()),
U16x3(typed_view) => DynamicImageView::U16x3(typed_view.into()),
U16x4(typed_view) => DynamicImageView::U16x4(typed_view.into()),
I32(typed_view) => DynamicImageView::I32(typed_view.into()),
F32(typed_view) => DynamicImageView::F32(typed_view.into()),
}
}
}
+21 -9
View File
@@ -1,18 +1,21 @@
use thiserror::Error;
#[derive(Error, Debug, Clone, Copy, PartialEq, Eq)]
pub enum ImageRowsError {
#[error("Count of rows don't match to image height")]
InvalidRowsCount,
#[error("Size of row don't match to image width")]
InvalidRowSize,
#[non_exhaustive]
pub enum ImageError {
#[error("Pixel type of image is not supported")]
UnsupportedPixelType,
}
#[derive(Error, Debug, Clone, Copy)]
#[error("Size of slice with pixels is smaller than required")]
pub struct InvalidPixelsSliceSize;
#[derive(Error, Debug, Clone, Copy, PartialEq, Eq)]
pub enum ImageBufferError {
#[error("Size of buffer is smaller than required")]
InvalidBufferSize,
#[error("Alignment of buffer don't match to alignment of u32")]
#[error("Alignment of buffer don't match to alignment of required pixel type")]
InvalidBufferAlignment,
}
@@ -26,9 +29,16 @@ pub enum CropBoxError {
WidthOrHeightLessOrEqualToZero,
}
#[derive(Error, Debug, Clone, Copy)]
#[error("Type of pixels of the source image is not equal to pixel type of the destination image")]
pub struct DifferentTypesOfPixelsError;
#[derive(Error, Debug, Clone, Copy, PartialEq, Eq)]
#[non_exhaustive]
pub enum ResizeError {
#[error("Source or destination image is not supported")]
ImageError(#[from] ImageError),
#[error("Pixel type of source image does not match to destination image")]
PixelTypesAreDifferent,
#[error("Source cropping option is invalid: {0}")]
SrcCroppingError(#[from] CropBoxError),
}
#[derive(Error, Debug, Clone, Copy)]
#[error(
@@ -38,6 +48,8 @@ pub struct DifferentDimensionsError;
#[derive(Error, Debug, Clone, Copy, PartialEq, Eq)]
pub enum MappingError {
#[error("Source or destination image is not supported")]
ImageError(#[from] ImageError),
#[error("The dimensions of the source image are not equal to the dimensions of the destination image")]
DifferentDimensions,
#[error("Unsupported combination of pixels of source and/or destination images")]
-214
View File
@@ -1,214 +0,0 @@
use std::num::NonZeroU32;
use crate::pixels::{PixelExt, PixelType};
use crate::{DynamicImageView, DynamicImageViewMut, ImageBufferError, ImageView, ImageViewMut};
#[derive(Debug)]
enum BufferContainer<'a> {
MutU8(&'a mut [u8]),
VecU8(Vec<u8>),
}
impl<'a> BufferContainer<'a> {
fn as_vec(&self) -> Vec<u8> {
match self {
Self::MutU8(slice) => slice.to_vec(),
Self::VecU8(vec) => vec.clone(),
}
}
}
/// Simple container of image data.
#[derive(Debug)]
pub struct Image<'a> {
width: NonZeroU32,
height: NonZeroU32,
buffer: BufferContainer<'a>,
pixel_type: PixelType,
}
impl<'a> Image<'a> {
/// Create empty image with given dimensions and pixel type.
pub fn new(width: NonZeroU32, height: NonZeroU32, pixel_type: PixelType) -> Self {
let pixels_count = (width.get() * height.get()) as usize;
let buffer = BufferContainer::VecU8(vec![0; pixels_count * pixel_type.size()]);
Self {
width,
height,
buffer,
pixel_type,
}
}
pub fn from_vec_u8(
width: NonZeroU32,
height: NonZeroU32,
buffer: Vec<u8>,
pixel_type: PixelType,
) -> Result<Self, ImageBufferError> {
let size = (width.get() * height.get()) as usize * pixel_type.size();
if buffer.len() < size {
return Err(ImageBufferError::InvalidBufferSize);
}
if !pixel_type.is_aligned(&buffer) {
return Err(ImageBufferError::InvalidBufferAlignment);
}
Ok(Self {
width,
height,
buffer: BufferContainer::VecU8(buffer),
pixel_type,
})
}
pub fn from_slice_u8(
width: NonZeroU32,
height: NonZeroU32,
buffer: &'a mut [u8],
pixel_type: PixelType,
) -> Result<Self, ImageBufferError> {
let size = (width.get() * height.get()) as usize * pixel_type.size();
if buffer.len() < size {
return Err(ImageBufferError::InvalidBufferSize);
}
if !pixel_type.is_aligned(buffer) {
return Err(ImageBufferError::InvalidBufferAlignment);
}
Ok(Self {
width,
height,
buffer: BufferContainer::MutU8(buffer),
pixel_type,
})
}
/// Creates a copy of the image.
pub fn copy(&self) -> Image<'static> {
Image {
width: self.width,
height: self.height,
buffer: BufferContainer::VecU8(self.buffer.as_vec()),
pixel_type: self.pixel_type,
}
}
#[inline(always)]
pub fn pixel_type(&self) -> PixelType {
self.pixel_type
}
#[inline(always)]
pub fn width(&self) -> NonZeroU32 {
self.width
}
#[inline(always)]
pub fn height(&self) -> NonZeroU32 {
self.height
}
/// Buffer with image pixels.
#[inline(always)]
pub fn buffer(&self) -> &[u8] {
match &self.buffer {
BufferContainer::MutU8(p) => p,
BufferContainer::VecU8(v) => v,
}
}
/// Mutable buffer with image pixels.
#[inline(always)]
pub fn buffer_mut(&mut self) -> &mut [u8] {
match &mut self.buffer {
BufferContainer::MutU8(p) => p,
BufferContainer::VecU8(ref mut v) => v.as_mut_slice(),
}
}
#[inline(always)]
pub fn into_vec(self) -> Vec<u8> {
match self.buffer {
BufferContainer::MutU8(p) => p.into(),
BufferContainer::VecU8(v) => v,
}
}
#[inline(always)]
pub fn view(&self) -> DynamicImageView {
macro_rules! get_dynamic_image {
($img_type: expr) => {
($img_type(ImageView::from_buffer(self.width, self.height, self.buffer()).unwrap()))
};
}
match self.pixel_type {
PixelType::U8 => get_dynamic_image!(DynamicImageView::U8),
PixelType::U8x2 => get_dynamic_image!(DynamicImageView::U8x2),
PixelType::U8x3 => get_dynamic_image!(DynamicImageView::U8x3),
PixelType::U8x4 => get_dynamic_image!(DynamicImageView::U8x4),
PixelType::U16 => get_dynamic_image!(DynamicImageView::U16),
PixelType::U16x2 => get_dynamic_image!(DynamicImageView::U16x2),
PixelType::U16x3 => get_dynamic_image!(DynamicImageView::U16x3),
PixelType::U16x4 => get_dynamic_image!(DynamicImageView::U16x4),
PixelType::I32 => get_dynamic_image!(DynamicImageView::I32),
PixelType::F32 => get_dynamic_image!(DynamicImageView::F32),
}
}
#[inline(always)]
pub fn view_mut(&mut self) -> DynamicImageViewMut {
macro_rules! get_dynamic_image {
($img_type: expr) => {
($img_type(
ImageViewMut::from_buffer(self.width, self.height, self.buffer_mut()).unwrap(),
))
};
}
match self.pixel_type {
PixelType::U8 => get_dynamic_image!(DynamicImageViewMut::U8),
PixelType::U8x2 => get_dynamic_image!(DynamicImageViewMut::U8x2),
PixelType::U8x3 => get_dynamic_image!(DynamicImageViewMut::U8x3),
PixelType::U8x4 => get_dynamic_image!(DynamicImageViewMut::U8x4),
PixelType::U16 => get_dynamic_image!(DynamicImageViewMut::U16),
PixelType::U16x2 => get_dynamic_image!(DynamicImageViewMut::U16x2),
PixelType::U16x3 => get_dynamic_image!(DynamicImageViewMut::U16x3),
PixelType::U16x4 => get_dynamic_image!(DynamicImageViewMut::U16x4),
PixelType::I32 => get_dynamic_image!(DynamicImageViewMut::I32),
PixelType::F32 => get_dynamic_image!(DynamicImageViewMut::F32),
}
}
}
/// Generic image container for internal purposes.
pub(crate) struct InnerImage<'a, P>
where
P: PixelExt,
{
width: NonZeroU32,
height: NonZeroU32,
pixels: &'a mut [P],
}
impl<'a, P> InnerImage<'a, P>
where
P: PixelExt,
{
pub fn new(width: NonZeroU32, height: NonZeroU32, pixels: &'a mut [P]) -> Self {
Self {
width,
height,
pixels,
}
}
#[inline(always)]
pub fn src_view(&self) -> ImageView<P> {
ImageView::from_pixels(self.width, self.height, self.pixels).unwrap()
}
#[inline(always)]
pub fn dst_view(&mut self) -> ImageViewMut<P> {
ImageViewMut::from_pixels(self.width, self.height, self.pixels).unwrap()
}
}
+86 -526
View File
@@ -1,554 +1,114 @@
use std::fmt::Debug;
use std::mem::ManuallyDrop;
use std::num::NonZeroU32;
use std::slice;
use crate::pixels::InnerPixel;
use crate::{ArrayChunks, ImageError, PixelType};
use crate::pixels::{GetCount, IntoPixelComponent, PixelComponent, PixelExt};
use crate::{CropBoxError, DifferentDimensionsError, ImageBufferError, ImageRowsError, PixelType};
/// A trait for getting access to image data.
pub trait ImageView {
type Pixel: InnerPixel;
/// A crop box parameters that may be used with [`ImageView`]
/// and [`DynamicImageView`](crate::DynamicImageView)
#[derive(Debug, Clone, Copy)]
pub struct CropBox {
pub left: f64,
pub top: f64,
pub width: f64,
pub height: f64,
}
/// Generic immutable image view.
#[derive(Debug, Clone)]
pub struct ImageView<'a, P>
where
P: PixelExt,
{
width: NonZeroU32,
height: NonZeroU32,
crop_box: CropBox,
rows: Vec<&'a [P]>,
}
impl<'a, P> ImageView<'a, P>
where
P: PixelExt,
{
pub fn new(
width: NonZeroU32,
height: NonZeroU32,
rows: Vec<&'a [P]>,
) -> Result<Self, ImageRowsError> {
check_rows_count_and_size(width, height, &rows)?;
Ok(Self {
width,
height,
crop_box: CropBox {
left: 0.,
top: 0.,
width: width.get() as _,
height: height.get() as _,
},
rows,
})
fn pixel_type(&self) -> PixelType {
Self::Pixel::pixel_type()
}
pub fn from_buffer(
width: NonZeroU32,
height: NonZeroU32,
buffer: &'a [u8],
) -> Result<Self, ImageBufferError> {
let size = (width.get() * height.get()) as usize * P::size();
if buffer.len() < size {
return Err(ImageBufferError::InvalidBufferSize);
}
let rows_count = height.get() as usize;
let pixels = align_buffer_to(buffer)?;
let rows = pixels
.chunks_exact(width.get() as usize)
.take(rows_count)
.collect();
Ok(Self {
width,
height,
crop_box: CropBox {
left: 0.,
top: 0.,
width: width.get() as _,
height: height.get() as _,
},
rows,
})
}
fn width(&self) -> u32;
pub fn from_pixels(
width: NonZeroU32,
height: NonZeroU32,
pixels: &'a [P],
) -> Result<Self, ImageBufferError> {
let size = (width.get() * height.get()) as usize;
if pixels.len() < size {
return Err(ImageBufferError::InvalidBufferSize);
}
let rows_count = height.get() as usize;
let rows = pixels
.chunks_exact(width.get() as usize)
.take(rows_count)
.collect();
Ok(Self {
width,
height,
crop_box: CropBox {
left: 0.,
top: 0.,
width: width.get() as _,
height: height.get() as _,
},
rows,
})
}
fn height(&self) -> u32;
pub fn pixel_type(&self) -> PixelType {
P::pixel_type()
}
pub fn width(&self) -> NonZeroU32 {
self.width
}
pub fn height(&self) -> NonZeroU32 {
self.height
}
pub fn crop_box(&self) -> CropBox {
self.crop_box
}
pub fn set_crop_box(&mut self, crop_box: CropBox) -> Result<(), CropBoxError> {
let width_f = self.width().get() as f64;
let height_f = self.height().get() as f64;
if crop_box.width <= 0. || crop_box.height <= 0. {
return Err(CropBoxError::WidthOrHeightLessOrEqualToZero);
}
if crop_box.left < 0.
|| crop_box.top < 0.
|| crop_box.left >= width_f
|| crop_box.top >= height_f
{
return Err(CropBoxError::PositionIsOutOfImageBoundaries);
}
let right = crop_box.left + crop_box.width;
let bottom = crop_box.top + crop_box.height;
if right > width_f || bottom > height_f {
return Err(CropBoxError::SizeIsOutOfImageBoundaries);
}
self.crop_box = crop_box;
Ok(())
}
/// Set a crop box to resize the source image into the
/// aspect ratio of destination image without distortions.
/// Returns iterator by slices with image rows.
///
/// `centering` used to control the cropping position. Use (0.5, 0.5) for
/// center cropping (e.g. if cropping the width, take 50% off
/// of the left side, and therefore 50% off the right side).
/// (0.0, 0.0) will crop from the top left corner (i.e. if
/// cropping the width, take all the crop off of the right
/// side, and if cropping the height, take all of it off the
/// bottom). (1.0, 0.0) will crop from the bottom left
/// corner, etc. (i.e. if cropping the width, take all the
/// crop off the left side, and if cropping the height take
/// none from the top, and therefore all off the bottom).
pub fn set_crop_box_to_fit_dst_size(
&mut self,
dst_width: NonZeroU32,
dst_height: NonZeroU32,
centering: Option<(f64, f64)>,
) {
// This function based on code of ImageOps.fit() from Pillow package.
// https://github.com/python-pillow/Pillow/blob/master/src/PIL/ImageOps.py
let centering = if let Some((x, y)) = centering {
(x.clamp(0.0, 1.0), y.clamp(0.0, 1.0))
} else {
(0.5, 0.5)
};
/// Note: An implementation must guaranty that all rows returned by iterator
/// have the same size and this size isn't less than the image width.
fn iter_rows(&self, start_row: u32) -> impl Iterator<Item = &[Self::Pixel]>;
// calculate aspect ratios
let width = self.width.get() as f64;
let height = self.height.get() as f64;
let image_ratio = width / height;
let required_ration = dst_width.get() as f64 / dst_height.get() as f64;
let crop_width;
let crop_height;
// figure out if the sides or top/bottom will be cropped off
if (image_ratio - required_ration).abs() < f64::EPSILON {
// The image is already the needed ratio
crop_width = width;
crop_height = height;
} else if image_ratio >= required_ration {
// The image is wider than what's needed, crop the sides
crop_width = required_ration * height;
crop_height = height;
} else {
// The image is taller than what's needed, crop the top and bottom
crop_width = width;
crop_height = width / required_ration;
}
let crop_left = (width - crop_width) * centering.0;
let crop_top = (height - crop_height) * centering.1;
self.set_crop_box(CropBox {
left: crop_left,
top: crop_top,
width: crop_width,
height: crop_height,
})
.unwrap();
}
#[inline(always)]
pub(crate) fn iter_4_rows<'s>(
&'s self,
/// Returns iterator by arrays with two image rows.
///
/// Note: An implementation must guaranty that all rows returned by iterator
/// have the same size and this size isn't less than the image width.
fn iter_2_rows(
&self,
start_y: u32,
max_y: u32,
) -> impl Iterator<Item = [&'a [P]; 4]> + 's {
let start_y = start_y as usize;
let max_y = max_y.min(self.height.get()) as usize;
let rows = self.rows.get(start_y..max_y).unwrap_or(&[]);
rows.chunks_exact(4).map(|rows| match *rows {
[r0, r1, r2, r3] => [r0, r1, r2, r3],
_ => unreachable!(),
})
max_rows: u32,
) -> ArrayChunks<impl Iterator<Item = &[Self::Pixel]>, 2> {
ArrayChunks::new(self.iter_rows(start_y).take(max_rows as usize))
}
#[inline(always)]
pub(crate) fn iter_2_rows<'s>(
&'s self,
/// Returns iterator by arrays with four image rows.
///
/// Note: An implementation must guaranty that all rows returned by iterator
/// have the same size and this size isn't less than the image width.
fn iter_4_rows(
&self,
start_y: u32,
max_y: u32,
) -> impl Iterator<Item = [&'a [P]; 2]> + 's {
let start_y = start_y as usize;
let max_y = max_y.min(self.height.get()) as usize;
let rows = self.rows.get(start_y..max_y).unwrap_or(&[]);
rows.chunks_exact(2).map(|rows| match *rows {
[r0, r1] => [r0, r1],
_ => unreachable!(),
})
max_rows: u32,
) -> ArrayChunks<impl Iterator<Item = &[Self::Pixel]>, 4> {
ArrayChunks::new(self.iter_rows(start_y).take(max_rows as usize))
}
#[inline(always)]
pub(crate) fn iter_rows<'s>(&'s self, start_y: u32) -> impl Iterator<Item = &'a [P]> + 's {
let start_y = start_y as usize;
let rows = self.rows.get(start_y..).unwrap_or(&[]);
rows.iter().copied()
}
#[inline(always)]
pub(crate) fn get_row(&self, y: u32) -> Option<&'a [P]> {
self.rows.get(y as usize).copied()
}
#[inline(always)]
pub(crate) fn iter_rows_with_step<'s>(
&'s self,
mut y: f64,
/// Returns iterator by image rows selected from image with given step.
///
/// Note: An implementation must guaranty that all rows returned by iterator
/// have the same size and this size isn't less than the image width.
fn iter_rows_with_step(
&self,
start_y: f64,
step: f64,
max_count: usize,
) -> impl Iterator<Item = &'a [P]> + 's {
let steps = (self.height.get() as f64 - y) / step;
let steps = (steps.max(0.).ceil() as usize).min(max_count);
(0..steps).map(move |_| {
// Safety of value of y guaranteed by calculation of steps count
let row = unsafe { *self.rows.get_unchecked(y as usize) };
max_rows: u32,
) -> impl Iterator<Item = &[Self::Pixel]> {
let steps = (self.height() as f64 - start_y) / step;
let steps = (steps.max(0.).ceil() as usize).min(max_rows as usize);
let mut rows = self.iter_rows(start_y as u32);
let mut y = start_y;
let mut next_row_y = start_y as usize;
let mut cur_row = None;
(0..steps).filter_map(move |_| {
let req_row_y = y as usize;
if next_row_y <= req_row_y {
for _ in next_row_y..=req_row_y {
cur_row = rows.next();
}
next_row_y = req_row_y + 1;
}
y += step;
row
})
}
#[inline(always)]
pub(crate) fn iter_cropped_rows<'s>(&'s self) -> impl Iterator<Item = &'a [P]> + 's {
let first_row = self.crop_box.top as usize;
let last_row = first_row + self.crop_box.height as usize;
let rows = unsafe { self.rows.get_unchecked(first_row..last_row) };
let first_col = self.crop_box.left as usize;
let last_col = first_col + self.crop_box.width as usize;
rows.iter()
// Safety guaranteed by method 'set_crop_box'
.map(move |row| unsafe { row.get_unchecked(first_col..last_col) })
}
}
/// Generic mutable image view.
#[derive(Debug)]
pub struct ImageViewMut<'a, P>
where
P: PixelExt,
{
width: NonZeroU32,
height: NonZeroU32,
rows: Vec<&'a mut [P]>,
}
impl<'a, P> ImageViewMut<'a, P>
where
P: PixelExt,
{
pub fn new(
width: NonZeroU32,
height: NonZeroU32,
rows: Vec<&'a mut [P]>,
) -> Result<Self, ImageRowsError> {
check_rows_count_and_size(width, height, &rows)?;
Ok(Self {
width,
height,
rows,
})
}
pub fn from_buffer(
width: NonZeroU32,
height: NonZeroU32,
buffer: &'a mut [u8],
) -> Result<Self, ImageBufferError> {
let size = (width.get() * height.get()) as usize * P::size();
if buffer.len() < size {
return Err(ImageBufferError::InvalidBufferSize);
}
let rows_count = height.get() as usize;
let pixels = align_buffer_to_mut(buffer)?;
let rows = pixels
.chunks_exact_mut(width.get() as usize)
.take(rows_count)
.collect();
Ok(Self {
width,
height,
rows,
})
}
pub fn from_pixels(
width: NonZeroU32,
height: NonZeroU32,
pixels: &'a mut [P],
) -> Result<Self, ImageBufferError> {
let size = (width.get() * height.get()) as usize;
if pixels.len() < size {
return Err(ImageBufferError::InvalidBufferSize);
}
let rows_count = height.get() as usize;
let rows = pixels
.chunks_exact_mut(width.get() as usize)
.take(rows_count)
.collect();
Ok(Self {
width,
height,
rows,
})
}
pub fn pixel_type(&self) -> PixelType {
P::pixel_type()
}
pub fn width(&self) -> NonZeroU32 {
self.width
}
pub fn height(&self) -> NonZeroU32 {
self.height
}
#[inline(always)]
pub(crate) fn iter_rows_mut(&mut self) -> slice::IterMut<&'a mut [P]> {
self.rows.iter_mut()
}
#[inline(always)]
pub(crate) fn iter_4_rows_mut<'s>(
&'s mut self,
) -> impl Iterator<Item = [&'s mut &'a mut [P]; 4]> {
self.rows.chunks_exact_mut(4).map(|rows| match rows {
[a, b, c, d] => [a, b, c, d],
_ => unreachable!(),
})
}
#[inline(always)]
pub(crate) fn get_row_mut<'s>(&'s mut self, y: u32) -> Option<&'s mut &'a mut [P]> {
self.rows.get_mut(y as usize)
}
/// Copy pixels from src_view.
pub(crate) fn copy_from_view(
&mut self,
src_view: &ImageView<P>,
) -> Result<(), DifferentDimensionsError> {
let src_crop_box = src_view.crop_box();
if src_crop_box.left != src_crop_box.left.round()
|| src_crop_box.top != src_crop_box.top.round()
|| src_crop_box.width != src_crop_box.width.round()
|| src_crop_box.height != src_crop_box.height.round()
{
// The crop box has fractional part in some his part
return Err(DifferentDimensionsError);
}
if self.width.get() != src_crop_box.width as u32
|| self.height.get() != src_crop_box.height as u32
{
return Err(DifferentDimensionsError);
}
self.rows
.iter_mut()
.zip(src_view.iter_cropped_rows())
.for_each(|(d, s)| d.copy_from_slice(s));
Ok(())
}
/// Create cropped version of the view.
pub fn crop(
self,
left: u32,
top: u32,
width: NonZeroU32,
height: NonZeroU32,
) -> Result<Self, CropBoxError> {
if left >= self.width.get() || top >= self.height.get() {
return Err(CropBoxError::PositionIsOutOfImageBoundaries);
}
let right = left + width.get();
let bottom = top + height.get();
if right > self.width.get() || bottom > self.height.get() {
return Err(CropBoxError::SizeIsOutOfImageBoundaries);
}
let row_range = (left as usize)..(right as usize);
let rows = self
.rows
.into_iter()
.skip(top as usize)
.take(height.get() as usize)
.map(|row| unsafe { row.get_unchecked_mut(row_range.clone()) })
.collect();
Ok(Self {
width,
height,
rows,
cur_row
})
}
}
impl<'a, P> From<ImageViewMut<'a, P>> for ImageView<'a, P>
where
P: PixelExt,
{
fn from(view: ImageViewMut<'a, P>) -> Self {
let rows = {
let mut old_rows = ManuallyDrop::new(view.rows);
let (ptr, length, capacity) =
(old_rows.as_mut_ptr(), old_rows.len(), old_rows.capacity());
unsafe { Vec::from_raw_parts(ptr as *mut &[P], length, capacity) }
};
ImageView {
width: view.width,
height: view.height,
crop_box: CropBox {
left: 0.,
top: 0.,
width: view.width.get() as _,
height: view.height.get() as _,
},
rows,
}
/// A trait for getting mutable access to image data.
pub trait ImageViewMut: ImageView {
/// Returns iterator by mutable slices with image rows.
///
/// Note: An implementation must guaranty that all rows returned by iterator
/// have the same size and this size isn't less than the image width.
fn iter_rows_mut(&mut self, start_row: u32) -> impl Iterator<Item = &mut [Self::Pixel]>;
/// Returns iterator by arrays with four mutable image rows.
///
/// Note: An implementation must guaranty that all rows returned by iterator
/// have the same size and this size isn't less than the image width.
fn iter_4_rows_mut(&mut self) -> ArrayChunks<impl Iterator<Item = &mut [Self::Pixel]>, 4> {
ArrayChunks::new(self.iter_rows_mut(0))
}
}
fn check_rows_count_and_size<T>(
width: NonZeroU32,
height: NonZeroU32,
rows: &[impl AsRef<[T]>],
) -> Result<(), ImageRowsError> {
if rows.len() != height.get() as usize {
return Err(ImageRowsError::InvalidRowsCount);
}
let row_size = width.get() as usize;
if rows.iter().any(|row| row.as_ref().len() != row_size) {
return Err(ImageRowsError::InvalidRowSize);
}
Ok(())
/// Conversion into an [ImageView].
pub trait IntoImageView {
/// Returns pixels type of the image if this type is supported by the crate.
fn pixel_type(&self) -> Option<PixelType>;
fn width(&self) -> u32;
fn height(&self) -> u32;
fn image_view<P: InnerPixel>(&self) -> Option<impl ImageView<Pixel = P>>;
}
fn align_buffer_to<T>(buffer: &[u8]) -> Result<&[T], ImageBufferError> {
let (head, pixels, _) = unsafe { buffer.align_to::<T>() };
if !head.is_empty() {
return Err(ImageBufferError::InvalidBufferAlignment);
}
Ok(pixels)
/// Conversion into an [ImageViewMut].
pub trait IntoImageViewMut: IntoImageView {
fn image_view_mut<P: InnerPixel>(&mut self) -> Option<impl ImageViewMut<Pixel = P>>;
}
fn align_buffer_to_mut<T>(buffer: &mut [u8]) -> Result<&mut [T], ImageBufferError> {
let (head, pixels, _) = unsafe { buffer.align_to_mut::<T>() };
if !head.is_empty() {
return Err(ImageBufferError::InvalidBufferAlignment);
}
Ok(pixels)
}
pub fn change_type_of_pixel_components<S, D, In, Out, CC>(
src_image: &ImageView<S>,
dst_image: &mut ImageViewMut<D>,
) -> Result<(), DifferentDimensionsError>
where
Out: PixelComponent,
In: IntoPixelComponent<Out>,
CC: GetCount,
S: PixelExt<Component = In, CountOfComponents = CC>,
D: PixelExt<Component = Out, CountOfComponents = CC>,
{
if src_image.width() != dst_image.width() || src_image.height() != dst_image.height() {
return Err(DifferentDimensionsError);
}
for (s_row, d_row) in src_image.rows.iter().zip(dst_image.rows.iter_mut()) {
let s_components = S::components(s_row);
let d_components = D::components_mut(d_row);
for (&s_comp, d_comp) in s_components.iter().zip(d_components) {
*d_comp = s_comp.into_component();
}
}
Ok(())
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn crop_view_mut() {
let mut image = crate::Image::new(
NonZeroU32::new(64).unwrap(),
NonZeroU32::new(32).unwrap(),
PixelType::U8,
);
let image_view: ImageViewMut<crate::pixels::U8> =
ImageViewMut::from_buffer(image.width(), image.height(), image.buffer_mut()).unwrap();
let cropped_view = image_view
.crop(
10,
10,
NonZeroU32::new(44).unwrap(),
NonZeroU32::new(12).unwrap(),
)
.unwrap();
assert_eq!(cropped_view.width().get(), 44);
assert_eq!(cropped_view.height().get(), 12);
assert_eq!(cropped_view.rows.len(), 12);
for row in cropped_view.rows.iter() {
assert_eq!(row.len(), 44);
}
}
/// Returns supported by the crate pixels type of the image or `ImageError` if the image
/// has not supported pixels type.
pub(crate) fn try_pixel_type(image: &impl IntoImageView) -> Result<PixelType, ImageError> {
image.pixel_type().ok_or(ImageError::UnsupportedPixelType)
}
+99
View File
@@ -0,0 +1,99 @@
use crate::{CropBoxError, ImageView, ImageViewMut};
fn check_crop_box(
image_view: &impl ImageView,
left: u32,
top: u32,
width: u32,
height: u32,
) -> Result<(), CropBoxError> {
let img_width = image_view.width();
let img_height = image_view.height();
if left >= img_width || top >= img_height {
return Err(CropBoxError::PositionIsOutOfImageBoundaries);
}
let right = left + width;
let bottom = top + height;
if right > img_width || bottom > img_height {
return Err(CropBoxError::SizeIsOutOfImageBoundaries);
}
Ok(())
}
macro_rules! cropped_image_impl {
($wrapper_name:ident<$view_trait:ident>, $doc:expr) => {
#[doc = $doc]
pub struct $wrapper_name<V: $view_trait + Sized> {
image_view: V,
left: u32,
top: u32,
width: u32,
height: u32,
}
impl<V: $view_trait + Sized> $wrapper_name<V> {
pub fn new(
image_view: V,
left: u32,
top: u32,
width: u32,
height: u32,
) -> Result<Self, CropBoxError> {
check_crop_box(&image_view, left, top, width, height)?;
Ok(Self {
image_view,
left,
top,
width,
height,
})
}
}
impl<V: $view_trait> ImageView for $wrapper_name<V> {
type Pixel = V::Pixel;
fn width(&self) -> u32 {
self.width
}
fn height(&self) -> u32 {
self.height
}
fn iter_rows(&self, start_row: u32) -> impl Iterator<Item = &[Self::Pixel]> {
let left = self.left as usize;
let right = left + self.width as usize;
self.image_view
.iter_rows(self.top + start_row)
.take((self.height - start_row) as usize)
// SAFETY: correct values of the left and the right
// are guaranteed by new() method.
.map(move |row| unsafe { row.get_unchecked(left..right) })
}
}
};
}
cropped_image_impl!(
CroppedImage<ImageView>,
"It is wrapper that provides [ImageView] for part of wrapped image."
);
cropped_image_impl!(
CroppedImageMut<ImageViewMut>,
"It is wrapper that provides [ImageViewMut] for part of wrapped image."
);
impl<V: ImageViewMut> ImageViewMut for CroppedImageMut<V> {
fn iter_rows_mut(&mut self, start_row: u32) -> impl Iterator<Item = &mut [Self::Pixel]> {
let left = self.left as usize;
let right = left + self.width as usize;
self.image_view
.iter_rows_mut(self.top + start_row)
.take((self.height - start_row) as usize)
// SAFETY: correct values of the left and the right
// are guaranteed by new() method.
.map(move |row| unsafe { row.get_unchecked_mut(left..right) })
}
}
+183
View File
@@ -0,0 +1,183 @@
use crate::images::{TypedImage, TypedImageMut};
use crate::pixels::InnerPixel;
use crate::{
ImageBufferError, ImageView, ImageViewMut, IntoImageView, IntoImageViewMut, PixelType,
};
#[derive(Debug)]
enum BufferContainer<'a> {
MutU8(&'a mut [u8]),
VecU8(Vec<u8>),
}
impl<'a> BufferContainer<'a> {
fn as_vec(&self) -> Vec<u8> {
match self {
Self::MutU8(slice) => slice.to_vec(),
Self::VecU8(vec) => vec.clone(),
}
}
}
/// Simple dynamic container of image data that provides [IntoImageView] and [IntoImageViewMut].
#[derive(Debug)]
pub struct Image<'a> {
width: u32,
height: u32,
buffer: BufferContainer<'a>,
pixel_type: PixelType,
}
impl Image<'static> {
/// Create an empty image with given dimensions and pixel type.
pub fn new(width: u32, height: u32, pixel_type: PixelType) -> Self {
let pixels_count = width as usize * height as usize;
let buffer = BufferContainer::VecU8(vec![0; pixels_count * pixel_type.size()]);
Self {
width,
height,
buffer,
pixel_type,
}
}
/// Create an image from vector with pixels data.
pub fn from_vec_u8(
width: u32,
height: u32,
buffer: Vec<u8>,
pixel_type: PixelType,
) -> Result<Self, ImageBufferError> {
let size = width as usize * height as usize * pixel_type.size();
if buffer.len() < size {
return Err(ImageBufferError::InvalidBufferSize);
}
if !pixel_type.is_aligned(&buffer) {
return Err(ImageBufferError::InvalidBufferAlignment);
}
Ok(Self {
width,
height,
buffer: BufferContainer::VecU8(buffer),
pixel_type,
})
}
}
impl<'a> Image<'a> {
/// Create an image with from slice with pixels data.
pub fn from_slice_u8(
width: u32,
height: u32,
buffer: &'a mut [u8],
pixel_type: PixelType,
) -> Result<Self, ImageBufferError> {
let size = width as usize * height as usize * pixel_type.size();
if buffer.len() < size {
return Err(ImageBufferError::InvalidBufferSize);
}
if !pixel_type.is_aligned(buffer) {
return Err(ImageBufferError::InvalidBufferAlignment);
}
Ok(Self {
width,
height,
buffer: BufferContainer::MutU8(buffer),
pixel_type,
})
}
#[inline]
pub fn pixel_type(&self) -> PixelType {
self.pixel_type
}
#[inline]
pub fn width(&self) -> u32 {
self.width
}
#[inline]
pub fn height(&self) -> u32 {
self.height
}
/// Buffer with image pixels data.
#[inline]
pub fn buffer(&self) -> &[u8] {
match &self.buffer {
BufferContainer::MutU8(p) => p,
BufferContainer::VecU8(v) => v,
}
}
/// Mutable buffer with image pixels data.
#[inline]
pub fn buffer_mut(&mut self) -> &mut [u8] {
match &mut self.buffer {
BufferContainer::MutU8(p) => p,
BufferContainer::VecU8(ref mut v) => v.as_mut_slice(),
}
}
#[inline]
pub fn into_vec(self) -> Vec<u8> {
match self.buffer {
BufferContainer::MutU8(p) => p.into(),
BufferContainer::VecU8(v) => v,
}
}
/// Creates a copy of the image.
pub fn copy(&self) -> Image<'static> {
Image {
width: self.width,
height: self.height,
buffer: BufferContainer::VecU8(self.buffer.as_vec()),
pixel_type: self.pixel_type,
}
}
/// Get typed version of the image.
pub fn typed_image<P: InnerPixel>(&self) -> Option<TypedImage<P>> {
if P::pixel_type() != self.pixel_type {
return None;
}
let typed_image = TypedImage::from_buffer(self.width, self.height, self.buffer()).unwrap();
Some(typed_image)
}
/// Get typed mutable version of the image.
pub fn typed_image_mut<P: InnerPixel>(&mut self) -> Option<TypedImageMut<P>> {
if P::pixel_type() != self.pixel_type {
return None;
}
let typed_image =
TypedImageMut::from_buffer(self.width, self.height, self.buffer_mut()).unwrap();
Some(typed_image)
}
}
impl<'a> IntoImageView for Image<'a> {
fn pixel_type(&self) -> Option<PixelType> {
Some(self.pixel_type)
}
fn width(&self) -> u32 {
self.width
}
fn height(&self) -> u32 {
self.height
}
fn image_view<P: InnerPixel>(&self) -> Option<impl ImageView<Pixel = P>> {
self.typed_image()
}
}
impl<'a> IntoImageViewMut for Image<'a> {
fn image_view_mut<P: InnerPixel>(&mut self) -> Option<impl ImageViewMut<Pixel = P>> {
self.typed_image_mut()
}
}
+73
View File
@@ -0,0 +1,73 @@
use std::ops::DerefMut;
use bytemuck::cast_slice_mut;
use image::DynamicImage;
use crate::image_view::try_pixel_type;
use crate::images::{TypedImage, TypedImageMut};
use crate::pixels::InnerPixel;
use crate::{ImageView, ImageViewMut, IntoImageView, IntoImageViewMut, PixelType};
impl IntoImageView for DynamicImage {
fn pixel_type(&self) -> Option<PixelType> {
match self {
DynamicImage::ImageLuma8(_) => Some(PixelType::U8),
DynamicImage::ImageLumaA8(_) => Some(PixelType::U8x2),
DynamicImage::ImageRgb8(_) => Some(PixelType::U8x3),
DynamicImage::ImageRgba8(_) => Some(PixelType::U8x4),
DynamicImage::ImageLuma16(_) => Some(PixelType::U16),
DynamicImage::ImageLumaA16(_) => Some(PixelType::U16x2),
DynamicImage::ImageRgb16(_) => Some(PixelType::U16x3),
DynamicImage::ImageRgba16(_) => Some(PixelType::U16x4),
_ => None,
}
}
fn width(&self) -> u32 {
self.width()
}
fn height(&self) -> u32 {
self.height()
}
fn image_view<P: InnerPixel>(&self) -> Option<impl ImageView<Pixel = P>> {
if let Ok(pixel_type) = try_pixel_type(self) {
if P::pixel_type() == pixel_type {
return TypedImage::<P>::from_buffer(self.width(), self.height(), self.as_bytes())
.ok();
}
}
None
}
}
impl IntoImageViewMut for DynamicImage {
fn image_view_mut<P: InnerPixel>(&mut self) -> Option<impl ImageViewMut<Pixel = P>> {
if let Ok(pixel_type) = try_pixel_type(self) {
if P::pixel_type() == pixel_type {
return TypedImageMut::<P>::from_buffer(
self.width(),
self.height(),
image_as_bytes_mut(self),
)
.ok();
}
}
None
}
}
fn image_as_bytes_mut(image: &mut DynamicImage) -> &mut [u8] {
match image {
DynamicImage::ImageLuma8(img) => (*img).deref_mut(),
DynamicImage::ImageLumaA8(img) => (*img).deref_mut(),
DynamicImage::ImageRgb8(img) => (*img).deref_mut(),
DynamicImage::ImageRgba8(img) => (*img).deref_mut(),
DynamicImage::ImageLuma16(img) => cast_slice_mut((*img).deref_mut()),
DynamicImage::ImageLumaA16(img) => cast_slice_mut((*img).deref_mut()),
DynamicImage::ImageRgb16(img) => cast_slice_mut((*img).deref_mut()),
DynamicImage::ImageRgba16(img) => cast_slice_mut((*img).deref_mut()),
_ => &mut [],
}
}
+11
View File
@@ -0,0 +1,11 @@
//! Contains different types of images and wrappers for them.
pub use cropped_image::*;
pub use dyn_image::*;
pub use typed_image::*;
mod cropped_image;
mod dyn_image;
mod typed_image;
#[cfg(feature = "image")]
mod image_crate;
+222
View File
@@ -0,0 +1,222 @@
use crate::pixels::InnerPixel;
use crate::{ImageBufferError, ImageView, ImageViewMut, InvalidPixelsSliceSize};
#[derive(Debug)]
enum PixelsContainer<'a, P> {
Borrowed(&'a mut [P]),
Owned(Vec<P>),
}
impl<'a, P: InnerPixel> PixelsContainer<'a, P> {
pub fn borrow(&self) -> &[P] {
match self {
PixelsContainer::Borrowed(p_ref) => p_ref,
PixelsContainer::Owned(vec) => vec,
}
}
pub fn borrow_mut(&mut self) -> &mut [P] {
match self {
PixelsContainer::Borrowed(p_ref) => p_ref,
PixelsContainer::Owned(vec) => vec,
}
}
}
/// Generic image container that provides [ImageView].
#[derive(Debug)]
pub struct TypedImage<'a, P> {
width: u32,
height: u32,
pixels: &'a [P],
}
impl<'a, P> TypedImage<'a, P> {
pub fn from_pixels(
width: u32,
height: u32,
pixels: &'a [P],
) -> Result<Self, InvalidPixelsSliceSize> {
let pixels_count = width as usize * height as usize;
if pixels.len() < pixels_count {
return Err(InvalidPixelsSliceSize);
}
Ok(Self {
width,
height,
pixels,
})
}
pub fn from_buffer(
width: u32,
height: u32,
buffer: &'a [u8],
) -> Result<Self, ImageBufferError> {
let pixels = align_buffer_to(buffer)?;
Self::from_pixels(width, height, pixels).map_err(|_| ImageBufferError::InvalidBufferSize)
}
}
impl<'a, P: InnerPixel> ImageView for TypedImage<'a, P> {
type Pixel = P;
fn width(&self) -> u32 {
self.width
}
fn height(&self) -> u32 {
self.height
}
fn iter_rows(&self, start_row: u32) -> impl Iterator<Item = &[Self::Pixel]> {
let width = self.width as usize;
let start = start_row as usize * width;
self.pixels
.get(start..)
.unwrap_or_default()
.chunks_exact(width)
}
fn iter_rows_with_step(
&self,
start_y: f64,
step: f64,
max_rows: u32,
) -> impl Iterator<Item = &[Self::Pixel]> {
let row_size = self.width as usize;
let steps = (self.height() as f64 - start_y) / step;
let steps = (steps.max(0.).ceil() as u32).min(max_rows);
let mut y = start_y;
let mut next_row_y = start_y as usize;
let mut cur_row = None;
(0..steps).filter_map(move |_| {
let cur_row_y = y as usize;
if next_row_y <= cur_row_y {
let start = cur_row_y * row_size;
let end = start + row_size;
cur_row = self.pixels.get(start..end);
next_row_y = cur_row_y + 1;
}
y += step;
cur_row
})
}
}
/// Generic mutable image container that provides [ImageView] and [ImageViewMut].
#[derive(Debug)]
pub struct TypedImageMut<'a, P: Default + Copy> {
width: u32,
height: u32,
pixels: PixelsContainer<'a, P>,
}
impl<P: Default + Copy> TypedImageMut<'static, P> {
pub fn new(width: u32, height: u32) -> Self {
let pixels_count = width as usize * height as usize;
Self {
width,
height,
pixels: PixelsContainer::Owned(vec![P::default(); pixels_count]),
}
}
}
impl<'a, P: InnerPixel> TypedImageMut<'a, P> {
pub fn from_pixels(
width: u32,
height: u32,
pixels: &'a mut [P],
) -> Result<Self, InvalidPixelsSliceSize> {
let pixels_count = width as usize * height as usize;
if pixels.len() < pixels_count {
return Err(InvalidPixelsSliceSize);
}
Ok(Self {
width,
height,
pixels: PixelsContainer::Borrowed(pixels),
})
}
// pub fn from_components(
// width: u32,
// height: u32,
// components: &'a mut [P::Component],
// ) -> Result<Self, ImageBufferError> {
// let components_count = width as usize * height as usize * P::count_of_components();
// if components.len() < components_count {
// return Err(ImageBufferError::InvalidBufferSize);
// }
// let pixels = align_buffer_to_mut(components)?;
// Ok(Self {
// width,
// height,
// pixels: PixelsContainer::Borrowed(pixels),
// })
// }
pub fn from_buffer(
width: u32,
height: u32,
buffer: &'a mut [u8],
) -> Result<Self, ImageBufferError> {
let size = width as usize * height as usize * P::size();
if buffer.len() < size {
return Err(ImageBufferError::InvalidBufferSize);
}
let pixels = align_buffer_to_mut(buffer)?;
Self::from_pixels(width, height, pixels).map_err(|_| ImageBufferError::InvalidBufferSize)
}
}
impl<'a, P: InnerPixel> ImageView for TypedImageMut<'a, P> {
type Pixel = P;
fn width(&self) -> u32 {
self.width
}
fn height(&self) -> u32 {
self.height
}
fn iter_rows(&self, start_row: u32) -> impl Iterator<Item = &[Self::Pixel]> {
let width = self.width as usize;
let start = start_row as usize * width;
self.pixels
.borrow()
.get(start..)
.unwrap_or_default()
.chunks_exact(width)
}
}
impl<'a, P: InnerPixel> ImageViewMut for TypedImageMut<'a, P> {
fn iter_rows_mut(&mut self, start_row: u32) -> impl Iterator<Item = &mut [Self::Pixel]> {
let width = self.width as usize;
let start = start_row as usize * width;
self.pixels
.borrow_mut()
.get_mut(start..)
.unwrap_or_default()
.chunks_exact_mut(width)
}
}
pub(crate) fn align_buffer_to<T>(buffer: &[u8]) -> Result<&[T], ImageBufferError> {
let (head, pixels, _) = unsafe { buffer.align_to::<T>() };
if !head.is_empty() {
return Err(ImageBufferError::InvalidBufferAlignment);
}
Ok(pixels)
}
pub(crate) fn align_buffer_to_mut<T>(buffer: &mut [u8]) -> Result<&mut [T], ImageBufferError> {
let (head, pixels, _) = unsafe { buffer.align_to_mut::<T>() };
if !head.is_empty() {
return Err(ImageBufferError::InvalidBufferAlignment);
}
Ok(pixels)
}
+21 -8
View File
@@ -1,30 +1,37 @@
#![doc = include_str!("../README.md")]
//!
//! ## Feature flags
#![doc = document_features::document_features!()]
pub use alpha::errors::*;
pub use array_chunks::*;
pub use change_components_type::*;
pub use color::mappers::*;
pub use color::PixelComponentMapper;
pub use convolution::*;
pub use dynamic_image_view::{
change_type_of_pixel_components_dyn, DynamicImageView, DynamicImageViewMut,
};
pub use cpu_extensions::CpuExtensions;
pub use crop_box::*;
pub use errors::*;
pub use image_view::{change_type_of_pixel_components, CropBox, ImageView, ImageViewMut};
pub use image_view::*;
pub use mul_div::MulDiv;
pub use pixels::PixelType;
pub use resizer::{CpuExtensions, ResizeAlg, Resizer};
pub use resizer::{ResizeAlg, ResizeOptions, Resizer, SrcCropping};
pub use crate::image::Image;
use crate::alpha::AlphaMulDiv;
#[macro_use]
mod utils;
mod alpha;
mod array_chunks;
mod change_components_type;
mod color;
mod convolution;
mod dynamic_image_view;
mod cpu_extensions;
mod crop_box;
mod errors;
mod image;
mod image_view;
pub mod images;
mod mul_div;
#[cfg(target_arch = "aarch64")]
mod neon_utils;
@@ -36,3 +43,9 @@ mod simd_utils;
pub mod testing;
#[cfg(target_arch = "wasm32")]
mod wasm32_utils;
/// This trait must be used in your code instead of [InnerPixel](crate::pixels::InnerPixel).
#[allow(private_bounds)]
pub trait PixelTrait: Convolution + AlphaMulDiv {}
impl<P: Convolution + AlphaMulDiv> PixelTrait for P {}

Some files were not shown because too many files have changed in this diff Show More