This work moves beyond closed-set segmentation (Mask2Former) to open-set detection using SAM and Grounding DINO.This work moves beyond closed-set segmentation (Mask2Former) to open-set detection using SAM and Grounding DINO.

Foundation Models for 3D Scenes: DINOv2 vs. CLIP for Instance Differentiation

2025/12/11 02:00

Abstract and 1 Introduction

  1. Related Works

    2.1. Vision-and-Language Navigation

    2.2. Semantic Scene Understanding and Instance Segmentation

    2.3. 3D Scene Reconstruction

  2. Methodology

    3.1. Data Collection

    3.2. Open-set Semantic Information from Images

    3.3. Creating the Open-set 3D Representation

    3.4. Language-Guided Navigation

  3. Experiments

    4.1. Quantitative Evaluation

    4.2. Qualitative Results

  4. Conclusion and Future Work, Disclosure statement, and References

2.2. Semantic Scene Understanding and Instance Segmentation

f 3D scenes. This domain has been thoroughly explored using closed-set vocabulary methods, including our prior work [1], which utilizes Mask2Former [7] for image segmentation. Various studies [18, 19, 20] have adopted a similar approach to achieve object segmentation, resulting in a closed-set framework. While these methods are effective, they are constrained by the limitation of predefined object categories. Our approach employs SAM [21] to acquire segmentation masks for open-set detection. Moreover, our methodology, distinct from many existing techniques that depend heavily on extensive pre-training or fine-tuning, integrates these models to forge a more comprehensive and adaptable 3D scene representation. This emphasizes enhanced semantic understanding and spatial awareness.

\ To improve the semantic understanding of the objects detected within our images, we harness detailed feature representations using two foundational models: CLIP [9] and DINOv2 [10]. DINOv2, a Vision Transformer trained through self-supervision, recognises pixel-level correspondences between images and captures spatial hierarchies. Compared to CLIP, DINOv2 more effectively distinguishes between two distinct instances of the same object type, which poses challenges for CLIP.

\ It’s crucial to differentiate individual instances following the semantic identification of objects. Early methods employed a Region Proposal Network (RPN) to predict bounding boxes for these instances [22]. Alternatively, some strategies suggest a generalized architecture for managing panoptic segmentation [23]. In our preceding approach, we utilized the segmentation model Mask2Former [7], which employs an attention mechanism to isolate object-centric features. Recent research also tackles semantic scene understanding using open vocabularies [24], utilizing multi-view fusion and 3D convolutions to derive dense features from an open-vocabulary embedding space for precise semantic segmentation. Our current pipeline leverages Grounding DINO [25] to generate bounding boxes, which are then input into the Segment Anything Model (SAM) [21] to produce individual object masks, thus enabling instance segmentation within the scene.

\

:::info Authors:

(1) Laksh Nanwani, International Institute of Information Technology, Hyderabad, India; this author contributed equally to this work;

(2) Kumaraditya Gupta, International Institute of Information Technology, Hyderabad, India;

(3) Aditya Mathur, International Institute of Information Technology, Hyderabad, India; this author contributed equally to this work.

(4) Swayam Agrawal, International Institute of Information Technology, Hyderabad, India;

(5) A.H. Abdul Hafez, Hasan Kalyoncu University, Sahinbey, Gaziantep, Turkey;

(6) K. Madhava Krishna, International Institute of Information Technology, Hyderabad, India.

:::


:::info This paper is available on arxiv under CC by-SA 4.0 Deed (Attribution-Sharealike 4.0 International) license.

:::

\

Disclaimer: The articles reposted on this site are sourced from public platforms and are provided for informational purposes only. They do not necessarily reflect the views of MEXC. All rights remain with the original authors. If you believe any content infringes on third-party rights, please contact service@support.mexc.com for removal. MEXC makes no guarantees regarding the accuracy, completeness, or timeliness of the content and is not responsible for any actions taken based on the information provided. The content does not constitute financial, legal, or other professional advice, nor should it be considered a recommendation or endorsement by MEXC.

You May Also Like

Luxembourg adds Bitcoin to its wealth fund, but what does that mean for Europe?

Luxembourg adds Bitcoin to its wealth fund, but what does that mean for Europe?

The post Luxembourg adds Bitcoin to its wealth fund, but what does that mean for Europe? appeared on BitcoinEthereumNews.com. Key Takeaways Why does Luxembourg’s move matter? It’s the first Eurozone nation to include Bitcoin in a sovereign wealth fund. How does it fit into Europe’s bigger picture? The UK is opening crypto ETNs to retail investors, and the EU’s ESMA is expanding its oversight. Luxembourg has become the first Eurozone country to invest part of its sovereign wealth fund in Bitcoin. During the presentation of the 2026 Budget at the Chambre des Deputes, Finance Minister Gilles Roth confirmed that the Fonds Souverain Intergenerationnel du Luxembourg (FSIL) — the nation’s sovereign wealth fund — has allocated 1% of its portfolio to Bitcoin. Luxembourg’s Bitcoin play According to Bob Kieffer, Director of the Treasury, the decision reflects “the growing maturity of this new asset class” and “leadership in digital finance.” Under the FSIL’s revised investment policy, up to 15% of total assets can now be placed in alternative investments. This includes investments in private equity, real estate, and crypto assets. The Bitcoin exposure, roughly €8.5 million [around $9 million USD], is being made through ETFs to avoid custody and operational risks. Kieffer also acknowledged differing opinions about the move. He said,  “Some might argue that we’re committing too little too late; others will point out the volatility and speculative nature of the investment. Yet, given the FSIL’s mission, a 1% allocation strikes the right balance while sending a clear message about Bitcoin’s long-term potential.” A cautious, but symbolic shift The FSIL, created in 2014 to preserve wealth across generations, now manages roughly €850 million. The announcement also comes on the back of Luxembourg tightening its digital asset regulatory framework, while preparing to implement DAC8. This new move will expand tax and reporting standards for crypto service providers in 2026. If Bitcoin continues to gain acceptance among sovereign investors, Luxembourg’s decision could…
Share
BitcoinEthereumNews2025/10/10 02:02
XRP Fractal Signals $6–$7 Surge by November Amid DLT Disruption

XRP Fractal Signals $6–$7 Surge by November Amid DLT Disruption

The post XRP Fractal Signals $6–$7 Surge by November Amid DLT Disruption appeared on BitcoinEthereumNews.com. XRP Fractal Analysis Hints at $6–$7 Breakout by Mid-November According to renowned market analyst EGRAG CRYPTO, XRP may be on the verge of a significant price movement. In his latest analysis, he points to a fractal formation pattern that suggests XRP could reach the $6–$7 range by mid-November.  Source: EGRAG CRYPTO This projection has quickly caught the attention of traders and long-term investors, as XRP’s current price remains well below this target. Fractals, often used in technical analysis, are recurring chart patterns that can help predict future price action by identifying historical similarities in market behavior.  Therefore, EGRAG CRYPTO argues that XRP is currently mirroring a previous structure that led to a notable rally. If this fractal setup plays out as expected, it could mark one of the most significant price surges for the digital asset in recent years. If XRP reaches $6–$7 by mid-November, it would mark a major win for investors and a symbolic breakthrough for a token that has endured regulatory battles and market volatility, validating its resilience and cementing its relevance in the evolving digital finance ecosystem. Meanwhile, a recent cup-and-handle pattern signalled that XRP had the potential of soaring to $15 by year-end with the altcoin presently trading at $3.04 per CoinGecko data.  DLT-Based Solutions: How Ripple and Stellar are Redefining Cross-Border Banking According to crypto observer SMQKE, distributed ledger technology (DLT)-based solutions are increasingly challenging the traditional correspondent banking model.  For decades, cross-border payments have relied on a chain of intermediaries, often resulting in slow settlements, high costs, and limited transparency. But with the rise of blockchain networks such as Ripple and Stellar, the industry is experiencing a seismic shift. The correspondent banking model depends on trust and pre-funded accounts, locking up liquidity and exposing banks to counterparty risk.  Transactions often take days to…
Share
BitcoinEthereumNews2025/09/19 16:12