Skip to content

Implement the Concat CUDA kernel - #1333

Merged
Hector Li (HectorSVC) merged 3 commits into
masterfrom
cuda/concat
Jul 3, 2019
Merged

Implement the Concat CUDA kernel#1333
Hector Li (HectorSVC) merged 3 commits into
masterfrom
cuda/concat

Conversation

@HectorSVC

Copy link
Copy Markdown
Contributor

Implement the Concat CUDA kernel instead of using cudaMemCpy in a loop which is slow.
The Concat may have hundreds or thousands of inputs or even more in some models. Current implementation using CudaMemCpy in a loop is very slow. Implement the CUDA kernel code will improve the performance.

block_offset = block_index - range_left;
break;
}
}

@ke1337 Ke Deng (ke1337) Jul 2, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you want to optimize for hundreds of inputs, running this loop for every CUDA thread might still be very expensive. I think it might be better to set it up as a lookup table in CPU and reuse, so CUDA kernel here only need to deal with a simple lookup. #Resolved

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

will update it, same for split


In reply to: 299692803 [](ancestors = 299692803)

block_size_inside_axis_dim_div.d_ +
offset;

output_data[id] = reinterpret_cast<const T*>(input_ptr[input_index])[input_pos];

@jignparm jignparm Jul 2, 2019

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It seem like every thread writes into output_data[id] only 1 time -- so 1 thread implies 1 output index is populated. If the output tensor is large (i.e. larger than number of threads in the grid), how are the remaining indexes being populated? #Resolved

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

there's also a blocksPerGrid, check line 52


In reply to: 299716784 [](ancestors = 299716784)

@ke1337 Ke Deng (ke1337) left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:shipit:

@HectorSVC
Hector Li (HectorSVC) merged commit 2a6c69d into master Jul 3, 2019
@HectorSVC
Hector Li (HectorSVC) deleted the cuda/concat branch July 3, 2019 06:09
Dmitri Smirnov (yuslepukhin) pushed a commit that referenced this pull request Mar 17, 2026
## Describe your changes

## Checklist before requesting a review
- [ ] Add unit tests for this change.
- [ ] Make sure all tests can pass.
- [ ] Update documents if necessary.
- [ ] Lint and apply fixes to your code by running `lintrunner -a`
- [ ] Is this a user-facing change? If yes, give a description of this
change to be included in the release notes.
- [ ] Is this PR including examples changes? If yes, please remember to
update [example
documentation](https://github.com/microsoft/Olive/blob/main/docs/source/examples.md)
in a follow-up PR.

## (Optional) Issue link
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants