Perspectives and Prospects on Transformer Architecture for Cross-Modal Tasks with Language and Vision

Andrew Shin, Masato Ishii, Takuya Narihira

International Journal of Computer Vision · 2022 · 48 citations · 102 references

Concepts

References

102