经过修订的条件t-SNE：超越最近邻 (Revised Conditional t-SNE: Looking Beyond the Nearest Neighbors) - 专知论文

会员服务 ·

0

最近邻 · 近邻 · 高维 · 可扩展性 · 扩展性 ·

2023 年 4 月 11 日

Revised Conditional t-SNE: Looking Beyond the Nearest Neighbors

翻译：经过修订的条件t-SNE：超越最近邻

Edith Heiter,Bo Kang,Ruth Seurinck,Jefrey Lijffijt

from arxiv, 20 pages including supplement

Conditional t-SNE (ct-SNE) is a recent extension to t-SNE that allows removal of known cluster information from the embedding, to obtain a visualization revealing structure beyond label information. This is useful, for example, when one wants to factor out unwanted differences between a set of classes. We show that ct-SNE fails in many realistic settings, namely if the data is well clustered over the labels in the original high-dimensional space. We introduce a revised method by conditioning the high-dimensional similarities instead of the low-dimensional similarities and storing within- and across-label nearest neighbors separately. This also enables the use of recently proposed speedups for t-SNE, improving the scalability. From experiments on synthetic data, we find that our proposed method resolves the considered problems and improves the embedding quality. On real data containing batch effects, the expected improvement is not always there. We argue revised ct-SNE is preferable overall, given its improved scalability. The results also highlight new open questions, such as how to handle distance variations between clusters.

翻译：条件t-SNE(ct-SNE)是近期对t-SNE的扩展，允许从嵌入式中去除已知集群信息，从而获得超出标签信息的结构。例如，在想要消除一组类之间不需要的差异时很有用。我们发现，在许多现实设置中，如果在原始高维空间中数据在标签上呈良好的集群，则ct-SNE会失败。我们通过对高维相似性进行条件约束，同时分别存储同类别和异类别最近邻居，提出了修订方法。这也使得可以使用最近提出的t-SNE加速方法，提高可扩展性。从对合成数据的实验中，我们发现我们的提议方法解决了考虑的问题并改善了嵌入式的质量。在包含批次效应的真实数据上，预期的改进并不总是存在。我们认为，鉴于改进的可扩展性，修订的ct-SNE总体上还是更好的。这些结果也突出了新的开放性问题，例如如何处理集群之间的距离变化。

0

相关内容

最近邻

对比学习简述

专知会员服务

90+阅读 · 2021年6月29日

【InterSpeech2020】混合语音识别系统中的词汇扩展技术，Techniques for Vocabulary Expansion in Hybrid Speech Recognition Systems

【InterSpeech2020】混合语音识别系统中的词汇扩展技术，Techniques for Vocabulary Expansion in Hybrid Speech Recognition Systems

专知会员服务

17+阅读 · 2020年3月23日

【SIGMOD2020】知识图谱补全方法的现实再评价，Realistic Re-evaluation of Knowledge Graph Completion Methods: An Experimental Study

【SIGMOD2020】知识图谱补全方法的现实再评价，Realistic Re-evaluation of Knowledge Graph Completion Methods: An Experimental Study

专知会员服务

33+阅读 · 2020年3月23日

【MIT】对抗鲁棒性的流形正则化，Manifold Regularization for Adversarial Robustness

【MIT】对抗鲁棒性的流形正则化，Manifold Regularization for Adversarial Robustness

专知会员服务

28+阅读 · 2020年3月11日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新七篇图像分割相关论文—Attention U-Net、对抗结构匹配损失、卷积CRFs、对抗样本、弱监督分割

【论文推荐】最新七篇图像分割相关论文—Attention U-Net、对抗结构匹配损失、卷积CRFs、对抗样本、弱监督分割

专知

19+阅读 · 2018年5月31日

【论文推荐】最新六篇对抗自编码器相关论文—多尺度网络节点表示、生成对抗自编码、逆映射、Wasserstein、条件对抗、去噪

【论文推荐】最新六篇对抗自编码器相关论文—多尺度网络节点表示、生成对抗自编码、逆映射、Wasserstein、条件对抗、去噪

专知

20+阅读 · 2018年4月7日

【论文推荐】最新5篇图像分割相关论文—条件随机场和深度特征学习、移动端网络、长期视觉定位、主动学习、主动轮廓模型、生成对抗性网络

【论文推荐】最新5篇图像分割相关论文—条件随机场和深度特征学习、移动端网络、长期视觉定位、主动学习、主动轮廓模型、生成对抗性网络

专知

13+阅读 · 2018年1月23日

基于CASSINI卫星观测的土星辐射带粒子动力学过程研究

国家自然科学基金

0+阅读 · 2014年12月31日

自发参量下转换产生的偏振纠缠光子对长寿命量子存储

国家自然科学基金

0+阅读 · 2014年12月31日

测地流的动力学研究

国家自然科学基金

0+阅读 · 2013年12月31日

电子回旋加热条件下的托卡马克等离子体输运研究

国家自然科学基金

0+阅读 · 2013年12月31日

红闪（sprite）对地闪的非线性响应与干涉效应研究

国家自然科学基金

0+阅读 · 2012年12月31日

双能谱CT迭代重建算法研究

国家自然科学基金

1+阅读 · 2012年12月31日

四能级倒Y型原子-环形腔耦合系统中非线性效应及窄带明亮纠缠光束产生的研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于list-mode数据的快速SART真3D PET断层重建算法的研究

国家自然科学基金

0+阅读 · 2011年12月31日

溶液晶体生长成核研究

国家自然科学基金

0+阅读 · 2009年12月31日

场论与粒子物理中的量子纠缠与退相干

国家自然科学基金

0+阅读 · 2008年12月31日

A joint estimation approach for monotonic regression functions in general dimensions

Arxiv

0+阅读 · 2023年5月28日

DiME: Maximizing Mutual Information by a Difference of Matrix-Based Entropies

Arxiv

0+阅读 · 2023年5月26日

Visual Information Matters for ASR Error Correction

Arxiv

0+阅读 · 2023年5月26日

Score-balanced Loss for Multi-aspect Pronunciation Assessment

Arxiv

0+阅读 · 2023年5月26日

The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model

Arxiv

0+阅读 · 2023年5月26日

Dynamic Inter-treatment Information Sharing for Heterogeneous Treatment Effects Estimation

Arxiv

0+阅读 · 2023年5月25日

Evaluating and reducing the distance between synthetic and real speech distributions

Arxiv

0+阅读 · 2023年5月25日

Mixture-of-Expert Conformer for Streaming Multilingual ASR

Arxiv

0+阅读 · 2023年5月25日

Prompt Distribution Learning

Arxiv

14+阅读 · 2022年5月6日

Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems

Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems

Arxiv

11+阅读 · 2019年11月4日

VIP会员

文章信息

相关主题

相关VIP内容

对比学习简述

专知会员服务

90+阅读 · 2021年6月29日

【InterSpeech2020】混合语音识别系统中的词汇扩展技术，Techniques for Vocabulary Expansion in Hybrid Speech Recognition Systems

【InterSpeech2020】混合语音识别系统中的词汇扩展技术，Techniques for Vocabulary Expansion in Hybrid Speech Recognition Systems

专知会员服务

17+阅读 · 2020年3月23日

【SIGMOD2020】知识图谱补全方法的现实再评价，Realistic Re-evaluation of Knowledge Graph Completion Methods: An Experimental Study

【SIGMOD2020】知识图谱补全方法的现实再评价，Realistic Re-evaluation of Knowledge Graph Completion Methods: An Experimental Study

专知会员服务

33+阅读 · 2020年3月23日

【MIT】对抗鲁棒性的流形正则化，Manifold Regularization for Adversarial Robustness

【MIT】对抗鲁棒性的流形正则化，Manifold Regularization for Adversarial Robustness

专知会员服务

28+阅读 · 2020年3月11日

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

【新书】数字图像(影像)处理手第二版，2176pdf，Mathematical Methods in Imaging

专知会员服务

93+阅读 · 2020年2月12日

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

社交网络上议题社群的公共焦虑研究，中国人民大学新闻学院塔娜讲师，第八届全国社会媒体处理大会SMP2019

专知会员服务

15+阅读 · 2019年10月23日

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

Connections between Support Vector Machines, Wasserstein distance and gradient-penalty GANs

专知会员服务

36+阅读 · 2019年10月17日

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

Deep Learning Based Detection and Correction of Cardiac MR Motion Artefacts During Reconstruction for High-Quality Segmentation

专知会员服务

59+阅读 · 2019年10月17日

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

Keras François Chollet 《Deep Learning with Python 》, 386页pdf

专知会员服务

160+阅读 · 2019年10月12日

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

【加州大学伯克利分校博士论文】通过自我监督预测学习泛化

专知会员服务

65+阅读 · 2019年10月9日

热门VIP内容

开通专知VIP会员享更多权益服务

人工智能治理的未来

模态感知的特征匹配：单一模态与跨模态技术的全面综述

无监督行人重识别研究综述

【牛津博士论文】面向神经影像应用的可扩展且可解释的空间模型

相关资讯

RoBERTa中文预训练模型：RoBERTa for Chinese

RoBERTa中文预训练模型：RoBERTa for Chinese

PaperWeekly

57+阅读 · 2019年9月16日

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

RoBERTa for Chinese：大规模中文预训练RoBERTa模型

AINLP

30+阅读 · 2019年9月8日

Hierarchically Structured Meta-learning

Hierarchically Structured Meta-learning

CreateAMind

27+阅读 · 2019年5月22日

Transferring Knowledge across Learning Processes

Transferring Knowledge across Learning Processes

CreateAMind

29+阅读 · 2019年5月18日

强化学习的Unsupervised Meta-Learning

强化学习的Unsupervised Meta-Learning

CreateAMind

18+阅读 · 2019年1月7日

Unsupervised Learning via Meta-Learning

Unsupervised Learning via Meta-Learning

CreateAMind

43+阅读 · 2019年1月3日

A Technical Overview of AI & ML in 2018 & Trends for 2019

A Technical Overview of AI & ML in 2018 & Trends for 2019

待字闺中

18+阅读 · 2018年12月24日

【论文推荐】最新七篇图像分割相关论文—Attention U-Net、对抗结构匹配损失、卷积CRFs、对抗样本、弱监督分割

【论文推荐】最新七篇图像分割相关论文—Attention U-Net、对抗结构匹配损失、卷积CRFs、对抗样本、弱监督分割

专知

19+阅读 · 2018年5月31日

【论文推荐】最新六篇对抗自编码器相关论文—多尺度网络节点表示、生成对抗自编码、逆映射、Wasserstein、条件对抗、去噪

【论文推荐】最新六篇对抗自编码器相关论文—多尺度网络节点表示、生成对抗自编码、逆映射、Wasserstein、条件对抗、去噪

专知

20+阅读 · 2018年4月7日

【论文推荐】最新5篇图像分割相关论文—条件随机场和深度特征学习、移动端网络、长期视觉定位、主动学习、主动轮廓模型、生成对抗性网络

【论文推荐】最新5篇图像分割相关论文—条件随机场和深度特征学习、移动端网络、长期视觉定位、主动学习、主动轮廓模型、生成对抗性网络

专知

13+阅读 · 2018年1月23日

相关论文

A joint estimation approach for monotonic regression functions in general dimensions

Arxiv

0+阅读 · 2023年5月28日

DiME: Maximizing Mutual Information by a Difference of Matrix-Based Entropies

Arxiv

0+阅读 · 2023年5月26日

Visual Information Matters for ASR Error Correction

Arxiv

0+阅读 · 2023年5月26日

Score-balanced Loss for Multi-aspect Pronunciation Assessment

Arxiv

0+阅读 · 2023年5月26日

The Curious Price of Distributional Robustness in Reinforcement Learning with a Generative Model

Arxiv

0+阅读 · 2023年5月26日

Dynamic Inter-treatment Information Sharing for Heterogeneous Treatment Effects Estimation

Arxiv

0+阅读 · 2023年5月25日

Evaluating and reducing the distance between synthetic and real speech distributions

Arxiv

0+阅读 · 2023年5月25日

Mixture-of-Expert Conformer for Streaming Multilingual ASR

Arxiv

0+阅读 · 2023年5月25日

Prompt Distribution Learning

Arxiv

14+阅读 · 2022年5月6日

Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems

Predictive Engagement: An Efficient Metric For Automatic Evaluation of Open-Domain Dialogue Systems

Arxiv

11+阅读 · 2019年11月4日

相关基金

基于CASSINI卫星观测的土星辐射带粒子动力学过程研究

国家自然科学基金

0+阅读 · 2014年12月31日

自发参量下转换产生的偏振纠缠光子对长寿命量子存储

国家自然科学基金

0+阅读 · 2014年12月31日

测地流的动力学研究

国家自然科学基金

0+阅读 · 2013年12月31日

电子回旋加热条件下的托卡马克等离子体输运研究

国家自然科学基金

0+阅读 · 2013年12月31日

红闪（sprite）对地闪的非线性响应与干涉效应研究

国家自然科学基金

0+阅读 · 2012年12月31日

双能谱CT迭代重建算法研究

国家自然科学基金

1+阅读 · 2012年12月31日

四能级倒Y型原子-环形腔耦合系统中非线性效应及窄带明亮纠缠光束产生的研究

国家自然科学基金

0+阅读 · 2012年12月31日

基于list-mode数据的快速SART真3D PET断层重建算法的研究

国家自然科学基金

0+阅读 · 2011年12月31日

溶液晶体生长成核研究

国家自然科学基金

0+阅读 · 2009年12月31日

场论与粒子物理中的量子纠缠与退相干

国家自然科学基金

0+阅读 · 2008年12月31日

微信扫码咨询专知VIP会员