You mentioned that you have tried training DuoduoCLIP on a ViT-B/16-based CLIP model. Would you be able to share the training and evaluation image embeddings and text embeddings generated by the ViT-B/16 CLIP model? If you don’t have the ViT-B/16 embeddings available, could you please let me know how you obtained the text descriptions for all the Zero123 renderings? Thanks!
You mentioned that you have tried training DuoduoCLIP on a ViT-B/16-based CLIP model. Would you be able to share the training and evaluation image embeddings and text embeddings generated by the ViT-B/16 CLIP model? If you don’t have the ViT-B/16 embeddings available, could you please let me know how you obtained the text descriptions for all the Zero123 renderings? Thanks!